A from-scratch implementation of a scaled-down GPT-2 model in PyTorch, trained on the Snappfood dataset for sentiment-controlled Persian text generation.
-
Updated
Nov 2, 2025 - Python
A from-scratch implementation of a scaled-down GPT-2 model in PyTorch, trained on the Snappfood dataset for sentiment-controlled Persian text generation.
LLM pipeline: data→tokenizer→attention→GPT train/eval→instruction FT→sampling. Reproducible, clean configs, RTX-4060 defaults, ready for AMP/LoRA/DDP.
VEHANT Causal Temporal Action Detection System: State-of-the-art deep learning for real-time fight/collapse detection in videos. Features causal attention, motion tokenization, MediaPipe skeletons, uncertainty quantification, and multi-task learning (class + bbox + temporal). 95% accuracy, ONNX/Docker-ready, 25ms GPU inference. 🚀
A decoder-only Transformer implemented from scratch in PyTorch for character-level name generation.
From-scratch, first-principles implementations of every major attention mechanism — from vanilla dot-product attention to multi-head, causal, and modern LLM-scale variants (MLA, GQA, Flash Attention). Every notebook derives the math by hand, verifies it with loops before vectorizing, and builds up to a reusable PyTorch module.
Decoder-only Transformer language model built from scratch with PyTorch — 4 layers, 8 heads, character-level tokenizer, trained on Tiny Shakespeare (~6M params, loss 8.69→0.83 over 5 epochs on P100). Foundational educational project; see STATUS.md.
From-scratch attention mechanisms in NumPy: scaled dot-product, self/cross, multi-head, causal/local, MQA/GQA, additive attention, and verified gradients.
To associate your repository with the causal-attention topic, visit your repo's landing page and select "manage topics."