Size or Architecture? A Parameter-Matched, From-Scratch Comparison of a Transformer and the Dragon Hatchling

Yash Singh - IIIT Una

Two character-level language models, built from scratch and trained identically on TinyShakespeare: a GPT-2-style Transformer and BDH (the "Dragon Hatchling", arXiv:2509.26507), a brain-inspired architecture using linear attention over a sparse, positive neuron field. A first, naive comparison gave BDH a 30-0 sweep of a position-controlled LLM judge - but BDH had roughly twice the parameters, so the result was confounded by size. After shrinking BDH to match the Transformer's parameter budget (818K vs 803K, matched within 2%) and retraining both from scratch, BDH still wins on every axis: perplexity 4.26 vs 5.76, more real words, lower validation loss, and a 29-0 decisive judge sweep at 97% swap-consistency (p = 3.7e-9). The advantage survives parameter matching, so on this task it is architectural, not a size effect. The broader lesson is methodological: a 2x parameter gap alone can manufacture a clean-looking sweep, and only a parameter-matched rerun separates architecture from size. Includes the full reproducible harness, both trained checkpoints, all completions, and both orderings of every judge verdict.

Download PDF - LaTeX source + code

All papers