Multimodal Generative Systems โข Video Understanding โข World Models โข Conference Speaker
Building next-generation generative AI systems that understand and create multimodal content from text to motion to photorealistic video.
|
๐น Transformer-based architectures for textโvideo synthesis |
๐ธ Distributed training pipelines (PyTorch DDP) |
โโ Generative Video Systems โ โโ ๐ง Transformer-based architectures โ โโ ๐จ Diffusion model research โ โโ โฐ Temporal consistency modeling โ โโ ๐ฅ Text-to-motion-to-video synthesis โ โโ World Models โ โโ ๐ Environment dynamics learning โ โโ ๐ฎ Predictive modeling โ โโ ๐ฎ Temporal coherence โ โโ Multimodal AI โโ ๐ Cross-modal fusion โโ ๐๏ธ Vision-language models โโ ๐ต Audio-visual learning
โโ Distributed Training โ โโ โก PyTorch DDP (8-GPU clusters) โ โโ ๐ Large-scale dataset pipelines โ โโ ๐ฌ Systematic experimentation โ โโ ๐ Mixed precision training โ โโ Production ML โ โโ ๐ ONNX deployment โ โโ โก TensorRT optimization โ โโ ๐ฆ Model quantization โ โโ ๐ฏ Inference acceleration โ โโ Research Infrastructure โโ ๐ง Experiment tracking โโ ๐ Metric dashboards โโ ๐ ๏ธ Reproducible pipelines
| Domain | Technologies | Experience Level |
|---|---|---|
| ๐ง Generative Modeling | Diffusion โข GANs โข VAEs โข Transformers โข Temporal Consistency | โโโโโโโโโโโโ 95% |
| ๐ฌ Video & Motion | Pose Estimation โข Temporal Modeling โข Sequence Alignment โข Motion Synthesis | โโโโโโโโโโโโ 90% |
| ๐ Multimodal Systems | Cross-modal Fusion โข Vision-Language โข Audio-Visual Learning | โโโโโโโโโโโโ 85% |
| โก Distributed Training | PyTorch DDP โข 8-GPU Clusters โข Large-scale Pipelines โข Experiment Tracking | โโโโโโโโโโโโ 88% |
| ๐ Model Deployment | ONNX โข TensorRT โข Quantization โข Inference Optimization | โโโโโโโโโโโโ 82% |
Run | Keeping Models Reliable in Live Systems Tobacco Dock, London
Practical strategies for deploying and maintaining reliable AI systems in production, presented to delegates, AI leaders, researchers, and engineering leaders.
|
MSc Computer Vision, ๐ Dissertation: Cross-modal latent fusion for multimodal face generation |
2 Research Papers โจ Generative AI Architectures ๐ฌ Computer Vision Systems |
AI Encode London Winner โก Real-time AI prototype ๐ Production-ready |
MSc Dissertation: Latent fusion for multimodal synthesis
Tech: VAEs โข Cross-modal Fusion โข PyTorch
Status: Published Research
A hands-on portfolio focused on model training, systems tradeoffs, reproducibility, and publishing deployable artifacts.
Repository: Deep-Learning-using-PyTorch
Case studies: ML Engineering Portfolio
Model artifacts: Hugging Face ยท pynk17
| Experiment | What I worked on | Notebook / Model |
|---|---|---|
| Custom DDPM / UNet | Built diffusion models from a small UNet baseline to a larger 256 ร 256 UNet2D DDPM with attention, BF16, gradient accumulation and periodic sampling | Notebook ยท Model |
| Stable Diffusion systems comparison | Fine-tuned Stable Diffusion v1.5 across FP32, FP16, and FP16 + gradient checkpointing to compare precision and memory-oriented training configurations | Notebook ยท FP32 ยท FP16 ยท FP16 + GC |
| BLOOM-7B1 LoRA | Adapted a 7B-scale causal LM with PEFT / LoRA while training about 0.11% of total parameters, using FP16 and gradient checkpointing | Notebook ยท Adapter |
| PyTorch foundations | Worked through tensor fundamentals and deep learning building blocks used throughout the later experiments | Overview ยท Basics |
PyTorch ยท Diffusers ยท Transformers ยท PEFT / LoRA ยท Stable Diffusion ยท DDPM ยท Mixed Precision ยท Gradient Checkpointing ยท Hugging Face Hub ยท Safetensors ยท Experiment Documentation
I'm actively seeking opportunities to collaborate on cutting-edge research problems:
| Research Area | Specific Interests |
|---|---|
| ๐ World Models | Environment dynamics โข Predictive modeling โข Spatial reasoning |
| ๐ฌ Video Generation | Temporal coherence โข Motion synthesis โข Controllable generation |
| ๐ Multimodal AI | Cross-modal fusion โข Unified representations โข Audio-visual learning |
| โฐ Temporal Reasoning | Long-range dependencies โข Sequence modeling โข Event prediction |
| ๐ฎ Interactive Systems | Real-time generation โข Human-AI collaboration โข Embodied AI |
๐ London, United Kingdom
๐ MSc Computer Vision, Robotics & ML @ University of Surrey
๐ฌ Researching at the intersection of generative AI, video understanding, and world models