design-research-analysis is the analysis and interpretation layer
in the CMU Design Research Collective design-research ecosystem. It turns recurring
design-study records and exported experiment artifacts into reproducible
findings.
It provides typed, reusable workflows for sequence, language, embedding-map, and statistical analysis over recurring event logs.
- Coverage reports total line coverage for the default deterministic test suite; CI requires at least 95%.
- Examples Passing reports checked-in example scripts that execute successfully in the examples workflow.
- API in Examples reports curated top-level
__all__exports referenced by runnable examples.N/Nmeans every supported top-level export appears in at least one example, and CI requires 100%.
Run make coverage, make examples-test, and make examples-coverage to reproduce these checks locally.
This package centers on reproducible analysis workflows with a curated top-level API:
- Unified-table coercion, validation, and mapper-based derived columns
- Dataset profiling, schema checks, and codebook generation
- Sequence modeling (Markov chains, discrete HMM, Gaussian HMM)
- Language analysis (semantic convergence trajectories, topic modeling, sentiment scoring)
- Embedding maps (PCA, t-SNE, UMAP, PaCMAP, TriMap) with clustering, comparison, and trajectory-plotting helpers
- Statistical wrappers (group comparisons, OLS regression, mixed-effects models, nonparametrics, and power)
- Portable analysis-result records and evidence-linked paper contributions
- Runtime provenance capture for reproducibility manifests
- Deterministic, integrity-checked paper-draft bundles with narrow data selection
- Top-level artifact handoff helpers for experiment exports
- A thin CLI for deterministic pipeline runs
Requires Python 3.12+.
Maintainer workflows target Python 3.12 (.python-version).
For a VS Code path that starts from PyPI and then shows the repository example
workflow, see
VS Code Start.
Install from PyPI:
python -m pip install --upgrade pip
python -m pip install design-research-analysisThe base install supports unified-table validation and Markov analysis:
import design_research_analysis as dran
rows = [
{"timestamp": "2026-01-01T00:00:00Z", "session_id": "s1", "event_type": "ideate"},
{"timestamp": "2026-01-01T00:00:05Z", "session_id": "s1", "event_type": "refine"},
]
report = dran.validate_unified_table(rows)
model = dran.fit_markov_chain_from_table(rows)
print(report.is_valid, model.states)Common install profiles:
python -m pip install "design-research-analysis[seq]"
python -m pip install "design-research-analysis[lang,embeddings]"
python -m pip install "design-research-analysis[maps,embeddings]" # text-driven maps
python -m pip install "design-research-analysis[maps]" # numeric features
python -m pip install "design-research-analysis[stats,data]"
python -m pip install "design-research-analysis[all]"Unified-table coercion, validation, and derived-column helpers ship in the base
install, so there is no separate table extra.
For contributor workflows:
python -m venv .venv
source .venv/bin/activate
make dev
make testRun a compact end-to-end example:
PYTHONPATH=src python examples/basic_usage.pyFor dependency profiles and release-check guidance, see Dependencies and Extras.
The package installs a design-research-analysis CLI:
design-research-analysis validate-table --input data/events.csv --summary-json artifacts/validate.json
design-research-analysis run-sequence --input data/events.csv --summary-json artifacts/sequence.json --mode markov
design-research-analysis run-language --input data/events.csv --summary-json artifacts/language.json --trajectory-csv artifacts/language_trajectory.csv
design-research-analysis run-embedding-maps --input data/events.csv --summary-json artifacts/embedding_maps.json --map-csv artifacts/embedding_maps.csv
design-research-analysis run-stats --input data/events.csv --summary-json artifacts/stats.json --mode regression --x-columns x1,x2 --y-column yThe Python API can start from files too at the main ingestion points. For
example, coerce_unified_table("data/events.csv") uses the base install;
profile_dataframe("data/events.csv") requires the data extra.
Start with examples/README.md for runnable scripts across all analysis families.
See the published documentation for quickstart, workflow guidance, schema details, CLI reference, and API docs. Use the design-research umbrella documentation for canonical whole-stack orientation and cross-package examples.
Build docs locally with:
make docs-check
make docs-buildThe supported public surface is whatever is exported from design_research_analysis.__all__.
Selected primary entry points include:
- Package metadata:
__version__ - Artifact handoff helpers:
validate_experiment_events,build_condition_metric_table_from_artifacts,build_event_table_from_artifacts - Table contracts:
UnifiedTableConfig,UnifiedTableValidationReport,coerce_unified_table,derive_columns,validate_unified_table - Sequence:
fit_markov_chain_from_table,fit_discrete_hmm_from_table,fit_text_gaussian_hmm_from_table,decode_hmm, plotting helpers, and result types - Language:
compute_language_convergence,compute_semantic_distance_trajectory,fit_topic_model,score_sentiment - Embedding maps:
embed_records,build_embedding_map,cluster_embedding_map,compare_embedding_maps,plot_embedding_map,plot_embedding_map_grid - Statistics:
compare_groups,fit_regression,fit_mixed_effects,permutation_test,bootstrap_ci, power helpers - Dataset + runtime:
profile_dataframe,validate_dataframe,generate_codebook,capture_run_context,attach_provenance,write_run_manifest - Paper support:
build_analysis_result,write_analysis_result,load_analysis_result,collect_analysis_paper_contributions - Verified bundle:
create_research_bundle,verify_research_bundle,BUNDLE_SCHEMA_VERSION
Contribution workflow and validation gates are documented in CONTRIBUTING.md.