A Transformer trained from scratch on a 24,525-pair parallel corpus.
Training and evaluation notebook with recorded results. No hosted model.
Overview
An encoder-decoder Transformer built with PyTorch and trained from scratch to translate short conversational English sentences into Urdu, alongside tokenizer and back-translation experiments.
Problem
Urdu is right-to-left, morphologically rich, and low-resource compared with European languages, so off-the-shelf tokenization and small datasets both work against you.
Approach
Separate SentencePiece tokenizers for English and Urdu (3,000 tokens each), a 4-layer encoder and 4-layer decoder Transformer (d_model 512, 8 heads), gradient clipping, and greedy decoding. mBART-50 was used to generate back-translated data for augmentation.
Architecture
Select a component to see what it does. Blue packets show the direction data moves.
- English text to SentencePiece EN
- SentencePiece EN to Encoder x4
- Encoder x4 to Decoder x4 (memory)
- Decoder x4 to SentencePiece UR
- SentencePiece UR to Urdu text
What it does
- Subword tokenization trained per language with SentencePiece, compared with HF BPE tokenizers.
- Transformer encoder-decoder trained from scratch.
- Back-translation of 100 lines with facebook/mbart-large-50-many-to-many-mmt.
- sacreBLEU evaluation on held-out sentences.
Technical challenges
- Fine-tuning mBART-50 itself ran out of GPU memory on Colab, so it was used for back-translation only.
Recorded results
- sacreBLEU
- 20.41
- 50 held-out test sentences
- Source: MT_EngtoUrdu.ipynb, cell 9
- Training loss
- 4.16 to 0.73
- Epoch 1 to epoch 10
- Source: MT_EngtoUrdu.ipynb
Limitations
- The BLEU score is on a small 50-sentence sample.
- The encoder was trained with a causal mask, which limits how much source context it uses; fixing this is the first planned improvement.