Engraving of Jules Verne as a robot

Jules Verne Bot & the Writing Machine

Three character-level models trained on the Extraordinary Journeys — a recurrent network and two Transformers — writing side by side, in your browser.

RNN · 4.10M · ctx 120 Transformer · 4.04M · ctx 120 Transformer · 4.05M · ctx 256 123-character vocabulary no server · no framework

The Controls

RNN weights0%
Transformer weights · ctx 256waiting
Transformer weights · ctx 120on demand

The RNN loads first so the page is usable immediately, then the 256-character Transformer downloads in the background. Each model is about 8 MB of 16-bit weights, so the third — the Transformer that matches the RNN's own 120-character window — is fetched only when you ask for it.

Press Enter to generate. The model writes one character at a time, so it continues your text rather than answering it.

0.70
0.1 cautious1.0 adventurous1.6 wild

Divides the logits before sampling. Low values pick the safest next character; high values flatten the distribution and invite chaos.

400 chars

The novels are hard-wrapped and the models reproduce those line breaks mid-sentence. Switch this off to see the raw character stream.

Reproducible within the page. The browser uses its own generator, so a seed does not reproduce the notebook's exact characters.

From the Machine

Loading the RNN weights…

Side by Side

The same seed, temperature and sampler through each model. The first two have the same size and the same 120-character window, so what separates them is only recurrence against attention; the third adds a longer window on top. Running all three fetches the 120-character Transformer if it is not loaded yet.

RNN · ctx 120waiting
Not yet run.
Transformer · ctx 120waiting
Not yet run.
Transformer · ctx 256waiting
Not yet run.

What the Three Models Are

Diagram comparing the RNN and the Transformer: how each carries the past, their cost per character, and the measured numbers
The two architectures. The RNN compresses the past into one fixed hidden state; the Transformer keeps every past character's keys and values and attends to them directly. That is the whole difference — the tokenizer, the corpus and the parameter count are held equal.

Measured

RunParametersContextVal lossms / char
RNN · GRU 10244,095,8671201.1460.15
Transformer · ctx 1204,011,5201201.1025.8
Transformer · ctx 2564,046,3362561.0348.1

All three models are trained on the same ten novels, with the Project Gutenberg license stripped out, and measured on the same held-out 5% of every book. At equal size and an equal 120-character window, attention beats recurrence by 0.044 nats per character — small, but about 22 times the seed-to-seed variation over three seeds each. Widening the window to 256 characters buys a further 0.068, so context length matters more here than the architecture change does.

The Transformer block expanded: layer norm, causal self-attention, residual, layer norm, feed-forward, and the ring-buffer KV cache
Inside the Transformer. The KV cache is the attention equivalent of the RNN's hidden state: each new character costs one pass instead of replaying the whole window. It cannot wrap, though — the positions are learned and absolute, so when the window fills the cache is rebuilt from the newest characters, once every 64 of them.
The RNN: embedding, a 1024-unit GRU with its three gates, a dense layer and sampling
Inside the RNN. The hidden state is the memory — everything read so far compressed into 1024 numbers, carried forward one character at a time.