Open-source Python library for evaluating ML model reliability beyond accuracy — with calibration, failure, and fairness diagnostics for informed deployment decisions.
-
Updated
Jul 20, 2026 - Python
Open-source Python library for evaluating ML model reliability beyond accuracy — with calibration, failure, and fairness diagnostics for informed deployment decisions.
Tests whether ML models preserve their behavior after conversion, quantization, or other transformations.
The course equips developers with techniques to enhance the reliability of LLMs, focusing on evaluation, prompt engineering, and fine-tuning. Learn to systematically improve model accuracy through hands-on projects, including building a text-to-SQL agent and applying advanced fine-tuning methods.
A reproducible, data-centric benchmarking framework evaluating the robustness of tabular machine learning models under systematic feature shift using OpenML-CC18 datasets and automated feature engineering.
Hard Reasoning Benchmark filtered with disagreement scores
Capability Schema Spec defines a shared semantic language for world model evaluation. Standardize capability definition, observation, and verification across models and benchmarks. Not a benchmark—a shared language. Define • Observe • Verify
PromptGuard is a pragmatic, opinionated framework for establishing continuous integration for LLM behavior. It operates on a simple, verifiable principle: run the same prompts across multiple model configurations, compare outputs against defined expectations, and flag semantic regressions.
Reproducible research framework for evaluating vision-model reliability under controlled image degradation, with calibration, failure-detection and prediction-level analysis.
Selective prediction for fetal ultrasound plane classification by pairing a ResNet-18 with diffusion reconstruction uncertainty and softmax confidence for calibrated human-AI deferral. Includes an honest ablation of uncertainty methods.
Reference implementation of the Capability Schema Specification. Proves that world model capabilities can be defined, observed, and verified in practice — with real checkpoints, real simulators, and real scores. Define • Observe • Verify • Deliver
Portfolio for MLOps and applied AI systems work
Participant-disjoint WESAD stress-monitoring reliability framework with calibration, threshold, false-alert, robustness, and subject-specific failure audits.
A reproducible visual-attribute verification framework combining group-disjoint evaluation, audited LoRA controls, calibration analysis, and CI-backed evidence contracts.
Enterprise-style RAG reliability platform for MLOps docs: cited answers, evals, traces, FastAPI, Next.js.
fraud-detection machine-learning xgboost model-monitoring model-reliability uncertainty-estimation out-of-distribution-detection meta-learning explainable-ai shap fastapi streamlit docker python scikit-learn
Code and results for "Stable Discrimination Can Hide Reliability Failures in AI Decision Support Under Distribution Shift and Changing Target Definitions". Computers 15(9), 560 (2026). DOI: 10.3390/computers15090560
Multi-LLM consensus engine for automated code review, diff analysis, and risk scoring.
To associate your repository with the model-reliability topic, visit your repo's landing page and select "manage topics."