class AbelYagubyan:
def __init__(self):
self.current_role = "Ph.D. Student in Artificial Intelligence @ University of Georgia"
self.building = "LearnOS, an open-source agentic AI university"
self.location = "San Francisco, CA"
self.education = {
"phd": "Ph.D. Artificial Intelligence @ University of Georgia (2026 - present)",
"masters": "M.S. Computer Science @ Northwestern (Summa Cum Laude)",
"bachelors": "B.A. Computer Science & Applied Math @ UC Berkeley",
}
self.research = [
"LLM-as-a-judge reliability",
"LLM agent reproducibility",
"Evaluation science",
"Open-source LLM evaluation tooling",
]
def achievements(self):
return {
"accelerator": "Y Combinator, Spring 2026 batch",
"papers": "4 published (GroundLM @ EMNLP 2026, MNRAS 2022, ACM SIGCSE 2022)",
"citations": "30, incl. Meta Superintelligence Labs and Alibaba",
"peer_review": "196 reviews across 9 Elsevier AI journals",
"open_source": "Area Triager @ TruLens (Snowflake), contributor @ UK AISI inspect_evals",
"industry": "Senior Data Scientist @ C3.ai, Co-founder @ FibonAI (Berkeley SkyDeck)",
}| Year | Paper | Venue | Cited by |
|---|---|---|---|
| 2026 | The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation (sole author) | GroundLM 2026 @ EMNLP 2026, Archival Long Paper | 6 (incl. Meta Superintelligence Labs) |
| 2026 | How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines (sole author) | arXiv preprint | 5 (incl. Alibaba) |
| 2022 | The Lick Observatory Supernova Search follow-up program: photometry data release of 70 SESNe (Zheng, Stahl, Filippenko, et al.) | MNRAS, vol. 512 | 19 |
| 2022 | Embedding of Programming IDEs into Computer-Based Testing Software (Yagubyan, Garcia) | ACM SIGCSE '22 | 0 |
Peer review: 196 verified reviews across nine Elsevier journals, including Engineering Applications of Artificial Intelligence, Image and Vision Computing, Neural Networks, and Information Fusion. See my ORCID record.
| Project | Role | What |
|---|---|---|
| LearnOS | Creator & Maintainer | Self-hosted, open-source "AI university": specialized agents (curriculum, Socratic tutor, assessment, research, analytics) that build personalized roadmaps, teach in real time, and grade toward mastery. One OpenRouter key, any model, no accounts. MIT. |
| TruLens (Snowflake) | Area Triager, listed in MAINTAINERS.md | 13 merged PRs: OpenTelemetry streaming metrics (TTFT, throughput) and correctness fixes across the database layer, evaluation API, providers, and dashboard. |
| inspect_evals (UK AI Security Institute) | Contributor | Migrated shared metric helpers to the framework's aggregation API; credited in release v0.20.0. |
π LearnOS
Open-source, agentic AI university with personalized roadmaps, real-time tutoring, and mastery-based grading
JavaScriptNode.jsViteSQLiteOpenRouter
ποΈ Synthora
AI development platform that builds full web and mobile apps from natural language: frontend, backend, database, workflows, deployment
TypeScriptNo-code
𧬠EHRJEPA
JEPA-based self-supervised framework for predictive patient embeddings from structured EHR data
PythonPyTorchSelf-supervised learning
π TableSage
Turns any CSV into an instant EDA and lightweight-modeling workspace with interpretable explanations
PythonData + ML
ποΈ ArxivScribe
Discord and Slack bot that monitors arXiv and posts LLM-generated TLDRs of the most relevant new ML papers daily
PythonLLMBots
π€ EasyVoiceClone Β· πΌοΈ Prompt2Deck
Minimal voice-cloning pipeline from a few audio samples Β· Presentation decks (PPTX, Slides, PDF) from a topic or outline
JavaScriptPythonVoice AISlide generation
gantt
title Professional Journey
dateFormat YYYY-MM
axisFormat %Y
section Research & Education
Ph.D. AI, University of Georgia :active, 2026-08, 2031-05
M.S. CS, Northwestern :2022-09, 2023-06
LBNL Research Contributor :2022-06, 2023-04
B.A. CS & Applied Math, Berkeley :2018-08, 2022-05
section Industry
Y Combinator (Spring 2026) :2026-03, 2026-06
Senior Data Scientist, C3.ai :2024-05, 2026-03
FibonAI Co-founder :2023-06, 2024-02
Apple x UC Berkeley Intern :2021-05, 2021-08
I'm always interested in collaborating on:
- βοΈ Evaluation science: making LLM judges and agents measurably reliable
- π§© Open-source LLM evaluation and observability tooling
- π Agentic systems for education
- π‘ Early-stage AI startups
π Visit my portfolio to learn more about my work!
"A judge that flips a coin is not a judge. Measure it."




