ANLP Assignment 1: Transformers from Scratch & Ablation Study

This repository contains trained model checkpoints and evaluation outputs for ANLP Assignment 1 across five controlled architectural configurations:

  • C1 (Base): Sinusoidal Absolute PE + Multi-Head Attention (MHA) + Pre-LayerNorm + Scratch BPE
  • C2 (RoPE): Rotary Positional Embedding (RoPE) + MHA + Pre-LayerNorm + Scratch BPE
  • C3 (GQA): Sinusoidal Absolute PE + Grouped-Query Attention (GQA) + Pre-LayerNorm + Scratch BPE
  • C4 (RMSNorm): Sinusoidal Absolute PE + MHA + RMSNorm + Scratch BPE
  • C5 (BLT): Sinusoidal Absolute PE + MHA + Pre-LayerNorm + Token-Free Byte Latent Transformer

Experimental Results Summary


----------------------------------------------------------------------------------------------------------------
               ANLP ASSIGNMENT 1 — ABLATION STUDY RESULTS
----------------------------------------------------------------------------------------------------------------
Config   | Status     | Val Loss   | Seq Acc    | Bit Acc    | Levenshtein  | BLEU     | VRAM (GB)  | Time (s)  
----------------------------------------------------------------------------------------------------------------
C1       | SUCCESS    | 5.7553     | 0.0%       | 0.0%       | 561.8        | 0.00     | 0.46       | 3.1       
C5       | SUCCESS    | 4.3724     | 0.0%       | 20.9%      | 460.6        | N/A      | 0.24       | 2.8       
----------------------------------------------------------------------------------------------------------------
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support