Data Scientist / ML Engineer at XLSmart — customer value & experience management: churn modeling, LTV forecasting, data pipelines at telco scale.
Interested in reliable ML systems and agentic AI: how models behave under distribution shift, how agents stay grounded when operating autonomously over long horizons, and what it takes to deploy these systems responsibly.
Portfolio · Email · LinkedIn · HuggingFace · Medium
Ensemble ML for Signal Generation in Indonesian Equity Markets — working draft, not submitted; reports a negative result
- Walk-forward study on the IDX: LightGBM + XGBoost ensemble with purged and embargoed cross-validation, ATR-scaled triple-barrier labeling, average-uniqueness sample weights, and a meta-labeler that scales position size by signal confidence
- Result: the model does not beat buy-and-hold. Equal-weight over 16 IDX large caps, 2025 out-of-sample, net of IDX costs: signal model +0.12% (Sharpe 0.05) with the meta-labeler filter, +7.62% (0.88) without it, +4.73% (1.38) on technical features only — against +14.17% (0.74) for buy-and-hold and +20.71% (1.12) for the IHSG
- Walk-forward accuracy inside the training period is 0.341–0.361 against a majority-class baseline of 0.532 on three-class labels; the model is worse than always predicting the most frequent label
- The earlier 195.6% / 2.70 Sharpe claim is withdrawn. Four evaluation defects account for it: the final model was trained on rows inside the test window, early stopping used the fold's own test set, the scaler was fitted on the whole series, and the 5-day calendar embargo was shorter than the 10-bar label horizon with no purge
- Code, run manifests (data SHA-256s), per-fold metrics, trade ledgers and figures:
project-aurum/research/
Turning-Point Analysis — bull/bear market phase dating for IDX stocks
- Censored local-extrema algorithm in the style of Pagan & Sossounov (2003) — no arbitrary fixed lookback window; phases must clear minimum duration and amplitude thresholds
- Isolates the 2020 COVID crash as a distinct bear phase on BBRI, unsupervised; properties enforced by a test suite
Indonesian suicide-ideation text classification — undergraduate thesis + transformer follow-up
- FastText + LSTM (thesis) against a fine-tuned IndoBERT on the identical train/val split: positive-class F1 0.79 → 0.90
- Class-imbalance treatment (class weighting, ADASYN) reduced F1 in the thesis arm — reported rather than dropped
- Deployed on HuggingFace · write-up on Medium
Project Aurum — quantitative trading system for the Indonesian Stock Exchange (IDX), 4-person team
- Ensemble signal model: LightGBM + XGBoost with purged and embargoed walk-forward cross-validation; SHAP-based explainability on every signal
- FastAPI backend, React + TypeScript frontend, real-time Telegram alerts, PostgreSQL + Redis
- A committed, reproducible re-evaluation lives in
research/: it fixes four leakage defects in the original pipeline and finds that the model does not beat buy-and-hold. Details in the Research section above - Stack: Python, LightGBM, XGBoost, scikit-learn, statsmodels, LangGraph, FastAPI, React
Autonomous Agent Orchestrator — self-hosted autonomous coding agent running on a schedule
- Priority queue with RAG-based context retrieval (Qdrant hybrid dense + sparse) — each task is enriched with relevant knowledge before execution
- LLM-as-judge scoring of outputs; failure memory stores past errors and retrieves them to avoid repeating mistakes
- Reversibility classification before executing destructive operations; immutable policy governance
- Automatic branch + PR workflow for team repos; direct commit for solo repos
- Stack: Python, Claude API, Qdrant, Redis, PostgreSQL, systemd
project-helios — telco CVM analytics & ML lab, built on the IBM Telco Customer Churn dataset and synthetic usage data at ~1M-row scale
- Idempotent DuckDB warehouse pipeline with a data-quality gate that aborts on critical failures before touching downstream tables
- Two independently-calibrated risk models (churn, late payment) — avoids miscalibration from conflating distinct risk types; churn AUC 0.823, late-payment AUC 0.628
- Forward-looking label construction with a leakage check enforced by test
- LLM-generated report narrative with explicit graceful degradation (missing key, API error, bad JSON all fall back safely)
- Stack: Python, DuckDB, scikit-learn, Claude API, GitHub Actions
MLBB Draft — Mobile Legends: Bang Bang draft assistant
- Hero recommendations scored across counter-matchups (40%), team synergy (35%), and meta strength (25%)
- Role gap detection — filters candidates by unfilled team roles before scoring
- Stack: Next.js, TypeScript, Tailwind
speech-event — event study on BBRI stock reaction to CEO-speech news
- Market-model regression (BBRI return ~ IHSG return) flags abnormal-return days, joined against scraped news by publish date
- Local LLM (Ollama, llama3) scores article sentiment, topic, and summary — no API dependency
- Streamlit dashboard for interactive price, article, event-study, and sentiment views
- Stack: Python, statsmodels, yfinance, Streamlit, PostgreSQL, Ollama
churn-project — telco churn prediction with a causal-inference layer
- XGBoost churn classifier with SHAP explainability, plus IPTW propensity-score estimation of the causal effect of product adoption on churn, revenue, and CLTV
- Distinguishes correlation from causation — flags which adoption pushes are worth doing vs. which need product fixes first
- Stack: Python, XGBoost, SHAP, scikit-learn, Streamlit
brimo-sentiment — sentiment analysis on public BRImo Play Store reviews
- Local LLM (Ollama, llama3) extracts topic, sentiment, and explanation per review — no per-request API cost at corpus scale
- Playwright-based Twitter scraper as a secondary text source
- Stack: Python, google-play-scraper, Playwright, Ollama
LapScout — Indonesian laptop recommendation platform (private repo, team project)
- 800+ models scraped from Indonesian retailers via Puppeteer-Stealth + ScrapingBee fallback
- Three-layer medallion pipeline (Bronze → Silver → Gold): raw scrape → normalized specs → fact table with market percentiles and SOTA scores
- AI chat advisor with function calling — natural language queries translate to live database lookups
- Stack: Next.js, TypeScript, Tailwind, PostgreSQL, Node.js
ML systems engineering · LLM agents and orchestration · Applied forecasting and signal generation · Validity and leakage control in time-series ML
Data & cloud platforms BigQuery · Snowflake · Databricks · AWS Bedrock · Tencent WeData
Languages Python · JavaScript
Databases PostgreSQL · MySQL · DuckDB
ML/DS scikit-learn · LightGBM · XGBoost · CatBoost · SHAP · BERT · statsmodels · PySpark
Tools LangChain · LangGraph · FastAPI · Docker · Redis · GitLab CI
Contributing idle compute and data from a personal homelab server:
Volunteer computing (BOINC) — 247,311 credits across Einstein@Home (gravitational-wave and pulsar signal analysis), MilkyWay@Home (galaxy structure modelling), and World Community Grid (cancer, COVID, and clean-energy research).
Indonesian NLP — a weekly pipeline collecting Bahasa Indonesia text (Indonesian Wikipedia extracts plus Indonesian news sites). The public release is deliberately restricted to records that carry full attribution, so every published record is traceable to the exact source revision: currently 552 Wikipedia extracts (~39k tokens) released under CC BY-SA 4.0 on HuggingFace, growing each week. Scraped news text and records collected before provenance was recorded are excluded from the release rather than published without a licence — 2,221 news-derived and ~30,000 unattributed records stay local.
Privacy infrastructure — a Tor Snowflake proxy supporting censorship circumvention.


