A clean PyTorch implementation of the original Transformer model + A German -> English translation example
-
Updated
Jan 24, 2022 - Python
A clean PyTorch implementation of the original Transformer model + A German -> English translation example
Repository for a transformer I coded from scratch and trained on the tiny-shakespeare dataset.
Transformer-based Beat Saber track generator
elaborate transformer implementation + detailed explanation
A Comprehensive Implementation of Transformers Architecture from Scratch
A structured, hands-on journey to mastering Transformers — from attention and mathematical foundations to building, training, and understanding modern Transformer architectures with PyTorch.
Implement the "Attention Is All You Need" paper from scratch using PyTorch, focusing on building a sequence-to-sequence transformer architecture for translating text from English to Italian
A complete implementation of the "Attention Is All You Need" Transformer model from scratch using PyTorch. This project focuses on building and training a Transformer for neural machine translation (English-to-Italian) on the OpusBooks dataset.
A showcase repository documenting the complete journey of building a Transformer and Neural Machine Translation framework in pure C—from tensor operations to automatic differentiation, training infrastructure, and end-to-end machine translation.
Modular Python implementation of encoder-only, decoder-only and encoder-decoder transformer architectures from scratch, as detailed in Attention Is All You Need.
This repository contains my coursework (assignments & semester exams) for the Natural Language Processing course at IIIT Delhi in Winter 2025.
PyTorch Transformer for neural machine translation (NMT), inspired by Attention Is All You Need. German→English on OPUS Books: training, inference, and attention visualization.
Simple and Easy Implimentation of Transformer Architecture Introduced in the paper "Attention Is All You Need"
A Transformer encoder built from first principles, implementing the core architecture from mathematical foundations to working PyTorch code, including tokenization, embeddings, positional encoding, self-attention, multi-head attention, LayerNorm, feed-forward networks, training, and evaluation.
From-scratch ~100k-parameter LLaMA decoder (RMSNorm, RoPE, SwiGLU, GQA), plus an equal-parameter ablation of each choice, a train-short/test-long probe of RoPE against learned absolute positions, and a parameter-free sweep of RoPE's rotation base. CPU-only, multi-seed; every README number renders from a committed artifact and CI byte-compares it.
_build a Transformer model from scratch using Pytorch
This project aims to build a Transformer from scratch and create a basic translation system from Arabic to English.
Implementation of Transformer:"Attention Is All You Need" in Pytorch
An educational implementation of core Transformer architecture concepts built from scratch using Python. This project explores how modern NLP transformer models work internally by implementing attention mechanisms, embeddings, positional encoding, and next-word prediction logic step-by-step.
Collection of implementations from scratch (mostly ML)
To associate your repository with the transformer-from-scratch topic, visit your repo's landing page and select "manage topics."