Pretrain, SFT, then the DPO run that held

I pretrained TinyGPT on tiny-shakespeare, fine-tuned it to say one sentence and stop, then used DPO to prefer a fuller line. TinyGPT has 12 layers, 8 heads, 128-dimensional embeddings, the GPT-2 byte-pair vocabulary, and about 15 million parameters. The pretrained checkpoint is out_bpe/tinygpt_shakespeare.pt. DPO at the fine-tuning learning rate raised the metrics and broke the generated text. The run kept here cuts that rate by ten. Every comparison uses ROMEO:, MENENIUS:, JULIET:, and KING RICHARD III:. ...

September 27, 2026 · 12 min · 2507 words · Rajat Patel