Built & Trained a Transformer from Scratch in Pure PyTorch for English-to-Tamil Machine Translation [Math + Code Breakdown] [P]
Hi everyone! 👋
I built and trained the complete Transformer architecture from scratch using pure PyTorch (`torch.nn` primitives) based on the original "Attention Is All You Need" paper.
I trained the model on an English-to-Tamil parallel translation dataset (`gopi30/english-tamil` on Hugging Face) using dual NVIDIA T4 GPUs on Kaggle.
I wrote a detailed mathematical breakdown and step-by-step tutorial covering every equation, tensor shape transformation, and PyTorch block.
Full Blog Post: https://imrancoder786.github.io/blog-post.html?post=transformer-from-scratch
GitHub Repository:
https://github.com/imrancoder786/ML_FROM_SCRATCH/tree/main/Transformer_from_scratch
I’d love to hear your feedback, suggestions, or any questions on the code/math!
I’d love to hear your feedback, suggestions, or any questions on the code/math!
[link] [comments]
Want to read more?
Check out the full article on the original site