MSc student in Molecular Biotechnology and Genetics at Istanbul University, working at the intersection of machine learning, genomics, and research software engineering.
I build reproducible, provenance-tracked computational systems that turn biological data into decision-grade evidence for healthcare and diagnostics — with a particular focus on making the validation honest, not just the model accurate.
🧬 KANIT — MSc Thesis
An end-to-end ML platform predicting antimicrobial resistance directly from bacterial whole-genome sequencing, delivering a versioned, queryable AMR biomarker knowledge base.
Scale: 45 models · 6 ESKAPEE organisms · 14 antibiotic classes · mean lineage-aware CV ROC-AUC 0.842 Scope: ~15.7k lines of Python across 46 modules · 220 commits · sole author
Highlights
- Out-of-core XGBoost over sparse matrices with millions of features — a custom streaming
DataIter→QuantileDMatrix, plus binary histogram quantisation (max_bin=2) for a 128× training-memory reduction - Alignment-free features — unitigs from a compacted de Bruijn graph, no reference genome required
- Lineage-aware validation (PopPUNK clusters +
StratifiedGroupKFold) as the reported metric. One model scores 0.985 on a random split but 0.429 under lineage-aware CV — a live demonstration of why validation design decides whether a clinical ML result is real - 7-layer evidence stack — CPSS stability selection with a PFER bound, TreeSHAP, permutation nulls, pyseer LMM, CARD/NCBI BLAST reverse-translation, SNP allele checks, external concordance vs. AMRFinderPlus/ResFinder
- Reproducibility engineering — three pinned Apptainer containers with committed conda lockfiles; every run stamps git commit, seed, config hash and eight tool versions into the knowledge base
- Delivered as a product — 13-table SQLite KB with semantic schema versioning, a FastAPI REST API, and a Streamlit explorer
- Registry-driven architecture — adding an organism is a registry entry, not a code change
- Operated on a national HPC cluster (SLURM), 45/45 models completed with zero failures
Python · XGBoost · Optuna · scikit-learn · SciPy sparse · FastAPI · SQLite · Apptainer · SLURM · pytest · GitHub Actions
🧫 bsc_thesis_ED_ML_AMR — BSc Thesis
The origin of the above: a feasibility study establishing that k-mer frequency features can classify beta-lactam resistance in E. coli, benchmarking Random Forest, Gradient Boosting and SVM under GridSearchCV.
Superseded by ML_AMR_Prediction_v2, which replaced per-chunk incremental training with full-data streaming boosting, chunk-level splits with lineage-aware grouped CV, and test-set threshold tuning with leakage-free training-derived thresholds.
📰 GeneGist
An automated, schedulable ETL pipeline that monitors NCBI PubMed, generates structured bilingual (EN/TR) science summaries with an LLM, and archives them to a cloud PostgreSQL database.
Designed for unattended operation: primary-key deduplication that short-circuits before any LLM call, all-or-nothing writes, and exponential retry backoff across two independently rate-limited APIs.
Python · Biopython/Entrez · Google Gemini · Supabase (PostgreSQL)
A seven-phase statistical ML pipeline for clinical diabetes prediction, built to demonstrate the mathematics rather than the API.
Includes maximum-likelihood estimation with a hand-derived analytic gradient (OLS ≡ MLE verified to 3.16 × 10⁻¹⁶), batch gradient descent from scratch, empirical bias–variance decomposition, logistic Z-statistics and odds ratios, and BCa bootstrap confidence intervals.
- Machine learning for clinical diagnostics and biomarker discovery
- Out-of-core / memory-bounded computation on high-dimensional data
- Reproducibility, provenance and validation engineering for scientific ML
- Bioinformatics pipelines and workflow automation
- Research software that other people can actually run
Languages — Python · SQL · Bash
ML & Statistics — XGBoost · scikit-learn · Optuna · SHAP/TreeSHAP · stability selection · grouped cross-validation · bootstrap inference · hypothesis testing with FDR control
Data & Scientific Computing — pandas · NumPy · SciPy (sparse CSR) · matplotlib/seaborn
Engineering — Git · pytest · GitHub Actions CI · ruff/mypy · FastAPI · SQLite · Apptainer/Singularity · SLURM/HPC · Linux · YAML-driven configuration
Bioinformatics — WGS analysis · unitigs / de Bruijn graphs · PopPUNK population genomics · BLAST+ · pyseer · CheckM2/QUAST · CARD/ARO ontology · NCBI Entrez
- LinkedIn: linkedin.com/in/demirbaseren
- Email: eren0demirbas@gmail.com



