Skip to content
View demirbase's full-sized avatar

Highlights

  • Pro

Organizations

@iumobg

Block or report demirbase

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
demirbase/README.md

Hi, I'm Eren Demirbaş 👋

MSc student in Molecular Biotechnology and Genetics at Istanbul University, working at the intersection of machine learning, genomics, and research software engineering.

I build reproducible, provenance-tracked computational systems that turn biological data into decision-grade evidence for healthcare and diagnostics — with a particular focus on making the validation honest, not just the model accurate.


🔬 Featured Projects

🧬 KANIT — MSc Thesis

An end-to-end ML platform predicting antimicrobial resistance directly from bacterial whole-genome sequencing, delivering a versioned, queryable AMR biomarker knowledge base.

Scale: 45 models · 6 ESKAPEE organisms · 14 antibiotic classes · mean lineage-aware CV ROC-AUC 0.842 Scope: ~15.7k lines of Python across 46 modules · 220 commits · sole author

Highlights

  • Out-of-core XGBoost over sparse matrices with millions of features — a custom streaming DataIter → QuantileDMatrix, plus binary histogram quantisation (max_bin=2) for a 128× training-memory reduction
  • Alignment-free features — unitigs from a compacted de Bruijn graph, no reference genome required
  • Lineage-aware validation (PopPUNK clusters + StratifiedGroupKFold) as the reported metric. One model scores 0.985 on a random split but 0.429 under lineage-aware CV — a live demonstration of why validation design decides whether a clinical ML result is real
  • 7-layer evidence stack — CPSS stability selection with a PFER bound, TreeSHAP, permutation nulls, pyseer LMM, CARD/NCBI BLAST reverse-translation, SNP allele checks, external concordance vs. AMRFinderPlus/ResFinder
  • Reproducibility engineering — three pinned Apptainer containers with committed conda lockfiles; every run stamps git commit, seed, config hash and eight tool versions into the knowledge base
  • Delivered as a product — 13-table SQLite KB with semantic schema versioning, a FastAPI REST API, and a Streamlit explorer
  • Registry-driven architecture — adding an organism is a registry entry, not a code change
  • Operated on a national HPC cluster (SLURM), 45/45 models completed with zero failures

Python · XGBoost · Optuna · scikit-learn · SciPy sparse · FastAPI · SQLite · Apptainer · SLURM · pytest · GitHub Actions


🧫 bsc_thesis_ED_ML_AMR — BSc Thesis

The origin of the above: a feasibility study establishing that k-mer frequency features can classify beta-lactam resistance in E. coli, benchmarking Random Forest, Gradient Boosting and SVM under GridSearchCV.

Superseded by ML_AMR_Prediction_v2, which replaced per-chunk incremental training with full-data streaming boosting, chunk-level splits with lineage-aware grouped CV, and test-set threshold tuning with leakage-free training-derived thresholds.


An automated, schedulable ETL pipeline that monitors NCBI PubMed, generates structured bilingual (EN/TR) science summaries with an LLM, and archives them to a cloud PostgreSQL database.

Designed for unattended operation: primary-key deduplication that short-circuits before any LLM call, all-or-nothing writes, and exponential retry backoff across two independently rate-limited APIs.

Python · Biopython/Entrez · Google Gemini · Supabase (PostgreSQL)


A seven-phase statistical ML pipeline for clinical diabetes prediction, built to demonstrate the mathematics rather than the API.

Includes maximum-likelihood estimation with a hand-derived analytic gradient (OLS ≡ MLE verified to 3.16 × 10⁻¹⁶), batch gradient descent from scratch, empirical bias–variance decomposition, logistic Z-statistics and odds ratios, and BCa bootstrap confidence intervals.


💻 What I work on

  • Machine learning for clinical diagnostics and biomarker discovery
  • Out-of-core / memory-bounded computation on high-dimensional data
  • Reproducibility, provenance and validation engineering for scientific ML
  • Bioinformatics pipelines and workflow automation
  • Research software that other people can actually run

🛠 Tech Stack

Languages — Python · SQL · Bash

ML & Statistics — XGBoost · scikit-learn · Optuna · SHAP/TreeSHAP · stability selection · grouped cross-validation · bootstrap inference · hypothesis testing with FDR control

Data & Scientific Computing — pandas · NumPy · SciPy (sparse CSR) · matplotlib/seaborn

Engineering — Git · pytest · GitHub Actions CI · ruff/mypy · FastAPI · SQLite · Apptainer/Singularity · SLURM/HPC · Linux · YAML-driven configuration

Bioinformatics — WGS analysis · unitigs / de Bruijn graphs · PopPUNK population genomics · BLAST+ · pyseer · CheckM2/QUAST · CARD/ARO ontology · NCBI Entrez


📫 Connect

Pinned Loading

  1. iumobg/KANIT iumobg/KANIT Public

    KANIT: an evidence-graded knowledge base of antimicrobial resistance biomarkers in the ESKAPEE pathogens

    Python 1 1

  2. GeneGist GeneGist Public

    🧬 Automated AI pipeline that monitors NCBI PubMed, generates bilingual (EN/TR) science blog posts using Google Gemini 1.5, and archives metadata in Supabase.

    Python

  3. iumobg/bsc_thesis_ED_ML_AMR iumobg/bsc_thesis_ED_ML_AMR Public

    Python

  4. Pima-Indians-Diabetes-ML-Pipeline Pima-Indians-Diabetes-ML-Pipeline Public

    End-to-end 7-phase ML pipeline on the Pima Indians Diabetes dataset: preprocessing, OLS/MLE, gradient descent from scratch, classification (LR/KNN/GNB), bias-variance, cross-validation, bootstrap, …

    Python