- π« Senior majoring in Computer Science and Applied Mathematics & Statistics at the Honors College @ SBU
- π Passionate about machine learning, RAG systems, and AI engineering
- π½ Based in Queens, New York
Iβm driven, detail-oriented, and care deeply about the quality of my work, especially when it comes to research and machine learning projects.
Recently, I was a Data Science Engineering Intern at AT&T, where I built AI-powered network documentation tools using RAG, FastMCP, and Gradio.
I'm also grateful to have been part of Break Through Tech AI, completing ML coursework with Cornell Tech and building geospatial forecasting models with Snowflake. At the McKinnon-Rosati Lab, I lead a research group developing high-precision doublet detection machine learning pipelines for scRNA-seq data.
What We Did: Led a team of 6 researchers to design an end-to-end Python pipeline combining unsupervised Leiden clustering and XGBoost classification to identify doublets in single-cell RNA sequencing data across 16 benchmark datasets (168K cells).
Tools: Python, scikit-learn, XGBoost, Leiden Clustering, PCA, Jupyter Notebooks
Result: Matched field-standard performance (scDblFinder) on accuracy and precision while achieving a lower false positive rate (5.5% vs 5.9%).
What We Did: Engineered a geospatial ML pipeline processing 600K+ MTA, OpenStreetMap, and population records to forecast transit demand and healthcare accessibility gaps across Brooklyn ZIP codes under 5 population-growth scenarios.
Tools: Python, SQL, Snowflake, HistGradientBoostingRegressor, Pandas, PyDeck, Streamlit
Result: Achieved a cross-validation score of 0.739 and deployed an interactive dashboard with demand heatmaps and subway overlays for urban planning and site selection.
What We Did: Built an analytical software tool to evaluate voting pattern changes, racially polarized voting (RPV), and legislative redistricting impacts using Ecological Inference (EI) and Gingles non-linear regression models.
Tools: Python, PyEI, Spring Boot, MongoDB, JavaScript, React, GIS Data Processing
Result: Successfully modeled voting turnout and candidate preference metrics across demographic splits to quantify the structural impacts of the repeal of the Voting Rights Act on electoral maps.
What We Did: Built and trained a 4-layer Convolutional Neural Network (CNN) with batch normalization, dropout, and grid search hyperparameter tuning on CIFAR-10.
Tools: Python, TensorFlow, Keras
Result: Achieved ~81.5% training accuracy and ~80.5% testing accuracy with minimal overfitting.
What We Did: Benchmark comparison evaluating 6 scRNA-seq doublet detection methods (scDblFinder, Scrublet, COMPOSITE, DoubletDetection) across artificial doublet generation and re-analysis experiments.
Tools: Python, R, scDblFinder, Scrublet, COMPOSITE, DoubletDetection
Results: Quantified method biases and performance tradeoffs to inform single-cell benchmark tool selection.
