AI-Assisted Support Automation | CI/CD Reliability | Python
I build reliable support and operations workflows with Python, APIs, CI/CD, human approval, and evidence-driven escalation.
My background combines enterprise technical support, Linux infrastructure, incident ownership, and developer-platform troubleshooting. My public projects use synthetic data and standalone implementations so I can demonstrate product thinking and engineering practices without exposing employer or customer information.
A production-shaped SRE and developer-platform project for ephemeral GitHub Actions runner fleets:
- OpenTofu configuration for a temporary AWS EKS lab
- Kubernetes, Helm, and Actions Runner Controller configuration
- Queue, startup, cleanup, and maximum-age SLOs
- Prometheus metrics and alert rules
- Capacity, cost, and incident evidence
- Read-only MCP tools for fleet inspection and recommendations
- Deterministic burst, capacity-loss, image-failure, and API-degradation tests
Current status: validated local and infrastructure scaffold; cloud deployment evidence has not yet been captured.
Tested Python tools and synthetic case-study patterns for:
- Evidence-led incident triage and next diagnostic steps
- Engineering-ready escalation quality checks
- Clear separation of observations, hypotheses, and root cause
- Incident communication, postmortems, and operational documentation
A read-only Python CLI that converts Linux fleet inventory into repeatable operational checks and guarded server-decommission plans:
- Patching, backup, monitoring, encryption, and OS-support readiness checks
- Explicit validation, prioritized findings, and JSON output
- Safety gates for approvals, dependencies, recovery evidence, and production
- Synthetic infrastructure data, unit tests, and CI-friendly exit codes
Policy-gated, approval-bound, staged Ansible rollouts for Linux fleets, shown through an SSSD directory migration:
- Planner that runs a staging canary first and never splits HA pairs into one wave
- Policy gate for wave size, freeze windows, readiness, and staging-before-production
- Approvals bound to a plan digest, with no self-approval and in-order wave apply
- Ansible role with automatic rollback, tested by Molecule on Rocky Linux 9 and Ubuntu 24.04
A Python CLI for repeatable CI/CD security and reliability reviews:
- Least-privilege permissions and immutable action references
- Unsafe untrusted-input and privileged pull-request patterns
- Timeouts, concurrency, OIDC, and reusable-workflow boundaries
- Human-readable and JSON output for local checks and CI
A dependency-free Python CLI that checks whether a support escalation still has a clear path to a customer outcome:
- Technical and customer-impact ownership, update cadence, and OOO coverage
- Business-day engineering response targets by severity
- Inactivity closures that are mislabeled as customer recovery
- Reopened escalations without a fresh technical decision
A gh CLI extension for a fast, read-only first look at a GitHub Enterprise
Server support bundle: disk, memory and OOM, failed services, and proxy or
connectivity errors.
- Support Knowledge Hygiene: score knowledge assets for ownership, freshness, usage, and duplication risk.
- Code Scanning Merge-Protection Reproduction: a minimal reproduction of analysis selection across Git references.
- Make the problem testable. Start with immutable evidence, define the failure boundary, and distinguish facts from hypotheses.
- Keep people in control. Require explicit review before generated content changes a customer-facing workflow.
- Build for reuse. Convert recurring investigations into tested tools, small reproductions, and practical documentation.
- Measure reliability. Evaluate expected behavior, unsafe claims, latency, and cost rather than relying on a polished demonstration alone.
- Improve the handoff. Give engineering a focused question, the evidence needed to answer it, and a clear customer-impact statement.
- Communicate uncertainty honestly. A useful outcome can be a bounded next step, not a premature root-cause claim.
Python 路 FastAPI 路 OpenTofu 路 Kubernetes 路 Helm 路 Prometheus 路
MCP 路 SLOs 路 GitHub Actions 路 CI/CD 路 Linux 路 Docker 路 APIs 路
Incident Response
- Secure and Reliable GitHub Actions Workflows
- GitHub Enterprise Server Operational Readiness
- GitHub Advanced Security Rollout and Operations
- From Symptom to Engineering-Ready Escalation
- Closure Is Not Customer Recovery
- Synthetic Support Engineering Case Studies
These are personal, unofficial projects. Public examples use synthetic data and do not contain employer source code, customer information, support-case data, or internal documentation.



