Three character-level models trained on the Extraordinary Journeys — a recurrent network and two Transformers — writing side by side, in your browser.
The RNN loads first so the page is usable immediately, then the 256-character Transformer downloads in the background. Each model is about 8 MB of 16-bit weights, so the third — the Transformer that matches the RNN's own 120-character window — is fetched only when you ask for it.
Press Enter to generate. The model writes one character at a time, so it continues your text rather than answering it.
Divides the logits before sampling. Low values pick the safest next character; high values flatten the distribution and invite chaos.
The novels are hard-wrapped and the models reproduce those line breaks mid-sentence. Switch this off to see the raw character stream.
Reproducible within the page. The browser uses its own generator, so a seed does not reproduce the notebook's exact characters.
The same seed, temperature and sampler through each model. The first two have the same size and the same 120-character window, so what separates them is only recurrence against attention; the third adds a longer window on top. Running all three fetches the 120-character Transformer if it is not loaded yet.
| Run | Parameters | Context | Val loss | ms / char |
|---|---|---|---|---|
| RNN · GRU 1024 | 4,095,867 | 120 | 1.146 | 0.15 |
| Transformer · ctx 120 | 4,011,520 | 120 | 1.102 | 5.8 |
| Transformer · ctx 256 | 4,046,336 | 256 | 1.034 | 8.1 |
All three models are trained on the same ten novels, with the Project Gutenberg license stripped out, and measured on the same held-out 5% of every book. At equal size and an equal 120-character window, attention beats recurrence by 0.044 nats per character — small, but about 22 times the seed-to-seed variation over three seeds each. Widening the window to 256 characters buys a further 0.068, so context length matters more here than the architecture change does.