Senior AI/ML Engineer — I build and run production AI systems end to end: agents, LLM platforms, data 📍 Amsterdam, Netherlands · Remote (EU) · 🇳🇱 Dutch citizen (EU work authorization)
Senior, hands-on AI/ML engineer with 15+ years in production software. Independent consultant since 2021 for Klarna, Booking.com, ING and others, now through my own company, PrimAxiom. I work directly with product and business stakeholders, mentor engineers, and go deep on the stack below the API: fine-tuning, small models, evaluation, and inference.
🌐 berkgokden.com (CV) · 💼 LinkedIn · 🤗 Hugging Face · ✉️ berk@primaxiom.ai
- Built one of Klarna's internal AI platforms, used by 5,000+ employees, with agents for meetings, voice, feedback and search — meeting efficiency up ~40%, lower cost per meeting
- Build AI agents at Booking.com for the AI Trip Planner and connected-trip work
- Cut infrastructure cost 90%+ with an AI load-prediction optimizer at Vamp.io (acquired by CircleCI), where I led a 12-person engineering team
- Build and run my own open-source AI systems: ReasonGraph (memory for AI agents) and Assay (calibrated decision models, 149M–27B)
I take on AI projects through PrimAxiom: a short paid scoping step, then a four-week pilot, or a small custom model on your own servers. Contract roles are welcome too. Email berk@primaxiom.ai · primaxiom.ai
| Project | What it is |
|---|---|
| reasongraph | Graph memory for AI agents: extracts entities and cause→effect relations with small open models, answers "why" questions with a cited chain. Library on PyPI, MCP server, and a hosted service (memory.primaxiom.ai). Measured in six languages in reasongraph-bench, including what did not help. |
| assay | Answers many typed questions about a text in one forward pass of a language model, with calibrated probabilities and an "enough evidence?" score. Full training pipeline; six models on Hugging Face. |
| causal-span-model | Multilingual cause/effect span tagger (mDeBERTa) used by ReasonGraph. |
| llama-constrain | Custom llama.cpp sampler for strict constrained / structured generation (GGUF). |
| personalens | A code review, but for UX: Playwright and a vision LLM review a URL from each persona's point of view, with scores you can track and gate in CI. |
| synthspan | Synthetic labeled-data generator for NER: templates and gazetteers or local-LLM few-shot with structured output, plus augmentation. |
| veri | Vector search engine in Go for ML feature serving. |
| 🤗 Berk | Small multilingual extraction models (place extraction in 13 languages, causal tagging), with ONNX versions for CPU. |
Agents — multi-agent orchestration · agent memory · RAG · MCP · LangChain/LangGraph · Claude Agent SDK Models & training — OpenAI · Anthropic · open-weight LLMs · fine-tuning (LoRA/QLoRA) · distillation into small models · embeddings · Whisper Evaluation & LLMOps — agent and LLM evaluation · Arize/Phoenix · reproducible benchmarks · guardrails · PII redaction Calibration — calibrated confidence (temperature scaling, Brier, ECE) · conformal prediction Extraction & retrieval — multilingual NER · relation extraction · GLiNER · cross-encoder rerankers · knowledge graphs Inference — vLLM · llama.cpp/GGUF · ONNX · constrained decoding Data & platform — Python · Go · TypeScript · SQL · PySpark · Snowflake · Kafka · pgvector · Kubernetes · Docker · AWS · GCP
Machine Learning Engineer at Booking.com (consultant via PrimAxiom) · Senior AI Engineer (contract) at Klarna · earlier ING, VodafoneZiggo, DPG Media, Caspar AI, Engineering Lead at Vamp.io (acquired by CircleCI), SAP. Founder and director of PrimAxiom.
Full CV: berkgokden.com



