Text Generation with RNNs

The Jules Verne Bot

Introduction

In this chapter, we will explore how to build a character-level text generation model using Recurrent Neural Networks (RNNs), specifically inspired by the works of Jules Verne.

The “Jules Verne Bot” project will help to show us the fundamental concepts of sequence modeling and text generation using deep learning techniques, as a preview of how the modern LLMs work.

Project Overview:

  • Goal: Create an AI model that generates text in the style of Jules Verne
  • Architecture: RNN with GRU (Gated Recurrent Unit) layers
  • Approach: Character-level text prediction
  • Framework: TensorFlow/Keras
  • Platform: Google Colab with Tesla T4 GPU
  • Extension: a size-matched Transformer trained on the same books, to see what attention changes (see From RNN to Transformer)

What Are We Actually Building?

Imagine we could teach a computer to write like Jules Verne, the famous author of “Twenty Thousand Leagues Under the Sea” and “Around the World in Eighty Days.” That’s precisely what we’re doing with the Jules Verne Bot. This project creates an artificial intelligence system that learns the patterns, style, and vocabulary from Jules Verne’s novels, then generates new text that sounds like it could have come from his pen.

Think of it like this: if we read enough of someone’s writing, we start to recognize their style. We notice they use certain phrases, prefer specific sentence structures, or have favorite topics. Our neural network does something similar, but with mathematical precision. It analyzes millions of characters from Verne’s works and learns to predict what character should come next in any given sequence.

Neural Network Architectures Background

Before we dive into the technical details, let’s understand why we use neural networks for this task and why we chose the specific type we did.

The Human Brain Analogy

When you read a sentence like “The submarine descended into the dark…” your brain automatically starts predicting what might come next. Maybe “depths” or “ocean” or “waters.” Your brain does this because it has learned patterns from all the text you’ve ever read. Neural networks work similarly, but they learn these patterns through mathematical calculations rather than biological processes.

Recurrent Neural Networks (RNN)

Before diving into our RNN implementation, let’s understand where RNNs fit in the neural network ecosystem:

Key Neural Network Architectures:

  • MLP (Multi-Layer Perceptron): Basic feedforward networks for general tasks, for example, vibration analysis
  • CNN (Convolutional Neural Networks): Specialized for image processing as Image Classification tasks and spatial data
  • RNN (Recurrent Neural Networks): Designed for sequential data like text and time series
  • GAN (Generative Adversarial Networks): Two networks competing for realistic data generation, as images
  • Transformers (Attention Networks): Modern architecture using attention mechanisms, as in LLMs (Large Language Models, such as GPT)

We chose a Recurrent Neural Network (RNN) for this project because text has a crucial property: order matters tremendously. The sequence “The cat sat on the mat” means something completely different from “Mat the on sat cat the.” Regular neural networks process all inputs simultaneously, like looking at a photograph. But for text, we need a network that processes information sequentially, remembering what came before to understand what should go next.

In text generation, we aim to predict the most probable word to follow a sentence.

Think of reading a book. You don’t just look at all the words on a page simultaneously. You read word by word, sentence by sentence, and your understanding builds as you progress. Each new word is interpreted in the context of everything you’ve read before in that chapter. RNNs work the same way.

The Memory Problem and GRU Solution

Early RNNs had a significant problem: they couldn’t remember information for very long. Imagine trying to understand a story where you could only remember the last few words you read. You’d lose track of characters, plot points, and context very quickly.

This is where the Gated Recurrent Unit (GRU) comes in. Think of GRU as an improved memory system with two special abilities:

Reset Gate: This decides when to “forget” old information. If the story switches to a new scene or character, the reset gate helps the network forget irrelevant details from the previous context.

Update Gate: This decides how much new information to incorporate. When encountering important plot points or character names, the update gate helps the network remember these crucial details for longer.

It’s like having a smart note-taking system that automatically decides what’s worth remembering and what can be forgotten.

Why RNNs for Text Generation?

Recurrent Neural Networks are designed explicitly for sequential data processing. Key characteristics:

  • Memory: RNNs maintain an internal state (memory) to remember previous inputs
  • Sequential Processing: Process data one element at a time, making them ideal for text
  • Variable Length Input: Can handle sequences of different lengths
  • Parameter Sharing: Same weights applied across different time steps

RNN Architecture Flow:

Input Sequence: x(t-1) → x(t) → x(t+1) → ...
Hidden State:   h(t-1) → h(t) → h(t+1) → ...
Output:         o(t-1) → o(t) → o(t+1) → ...

Dataset Preparation

Our model is trained on a curated collection of 10 classic Jules Verne novels, downloaded from public domain texts of the Gutenberg Project:

  1. “A Journey to the Centre of the Earth”
  2. “In Search of the Castaways”
  3. “An Antarctic Mystery”
  4. “In the year 2889”
  5. “Around the World in Eighty Days”
  6. “Michael Strogoff”
  7. “Five Weeks in a Balloon”
  8. “The Mysterious Island”
  9. “From the Earth to the Moon”
  10. “Twenty Thousand Leagues under the Sea”

“A Journey to the Centre of the Earth” teaches the model about geological descriptions and underground adventures. “Twenty Thousand Leagues Under the Sea” provides vocabulary about marine life and submarine technology. “Around the World in Eighty Days” offers geographical references and travel descriptions. Each book contributes unique vocabulary and stylistic elements while maintaining Verne’s consistent voice.

The complete dataset contains 5,768,791 characters, of which 123 are unique. To put this in perspective, that’s roughly equivalent to 3,000 pages of text. This provides our neural network with ample material to learn from, enabling it to capture both common patterns and unique expressions in Verne’s writing.

Data Preprocessing Steps

The text is used as it is: no lowercasing, no removal of punctuation. Case and punctuation are part of Verne’s style, and the model has to learn them too.

# Read all the books and join them into a single string
path = './books/'
text = ""
for filename in os.listdir(path):
    if filename.endswith(".txt"):
        with open(os.path.join(path, filename), 'r', encoding='utf-8') as file:
            text += file.read() + "\n"   # a newline between books

print(f"Total characters: {len(text)}")   # 5,768,791

A detail that turned out to matter: each Project Gutenberg file carries a header and, at the end, the full Gutenberg license. That boilerplate is 3.4% of the text, and the model learns it as faithfully as it learns Verne. You will see it in the generated sample later in this chapter. The From RNN to Transformer section removes it before training.

Tokenization and Vocabulary

Character-Level Tokenization

Here’s where our approach differs from how humans typically think about text. While we naturally think in words and sentences, our model processes text character by character. This means it learns that certain letters frequently follow others, that spaces separate words, and that punctuation marks signal sentence boundaries.

Why choose character-level processing? Consider the word “extraordinary,” which appears frequently in Verne’s work. A word-level model would need to have seen this exact word during training to use it. But a character-level model can generate this word by learning that ‘e’ often starts words, ‘x’ can follow ‘e’, ‘t’ often follows ‘x’, and so on. This allows our model to create new words or handle misspellings gracefully.

The downside is that character-level processing requires more computational steps to generate the same amount of text. Generating “Hello world” requires 11 prediction steps instead of just 2. However, for our educational purposes, this trade-off provides valuable insights into how language generation works at its most fundamental level.

Unlike word-level tokenization, character-level tokenization treats each character as a token.

Advantages of Character-Level Tokenization:

  • No Out-of-Vocabulary Issues: Every possible character sequence can be generated
  • Smaller Vocabulary: Only 123 unique characters vs thousands of words
  • Language Agnostic: Works with any language or symbol system
  • Handles Rare Words: Can generate new words character by character

Please see the following site for a great general visual explanation, from Andrej Karpathy, The Unreasonable Effectiveness of Recurrent Neural Networks.

Vocabulary Building Process

Computers work with numbers, not letters, so we need to convert our text into a numerical representation. We start by finding every unique character in our dataset. This includes not just letters A-Z and a-z, but also numbers, punctuation marks, spaces, and even special characters that might appear in the original texts.

Our Jules Verne collection contains 123 unique characters. These include obvious ones like letters and common punctuation, but also less common characters like accented letters from French names or special typography marks from the original publications.

Creating the Character Dictionary

We create two dictionaries: one that converts characters to numbers (encoding) and another that converts numbers back to characters (decoding). For example:

‘a’ might become 47, ‘b’ becomes 48, ‘c’ becomes 49, and so on. The space character might be 1, and the period might be 72. These assignments are arbitrary but consistent throughout our project.

When we want to process the phrase “The sea”, we convert it to something like [84, 72, 69, 1, 83, 69, 47]. When the model generates numbers like [84, 72, 69, 1, 87, 47, 83], we convert them back to “The was” (as an example).

# Create character-to-index mapping
text = "Your complete dataset text here..."
vocab = sorted(set(text))
char_to_idx = {char: idx for idx, char in enumerate(vocab)}
idx_to_char = {idx: char for idx, char in enumerate(vocab)}

print(f"Vocabulary size: {len(vocab)}")
print(f"Unique characters: {vocab}")

It is possible to experiment with (sub-word level) tokenization using OpenAI’s tokenizer tool at:

https://platform.openai.com/tokenizer

Training Sequences

The Sliding Window Approach

Our model learns by playing a sophisticated prediction game. We show it sequences of 120 characters and ask it to predict what the 121st character should be. Think of it like a fill-in-the-blank exercise, but instead of missing words, we’re missing the next character.

In practice, the text is cut into consecutive, non-overlapping chunks of 121 characters. The first 120 are the input, and the same 120 shifted by one character are the target. So each chunk is not one example but 120: the model predicts the second character from the first, the third from the first two, and so on up to the 121st.

Training Configuration

  • Sequence Length: 120 characters (approximately one paragraph)
  • Input-Output Relationship: Predict the next character given the previous characters

Why 120 Characters?

We chose 120 characters as our context window because it roughly corresponds to one paragraph of text in English. This gives the model enough context to understand local patterns (like completing words and phrases) while remaining computationally manageable. In practical terms, 120 characters might look like:

“The Nautilus had been cruising in these waters for some time. Captain Nemo stood on the bridge, observing the vast exp”

From this context, the model might predict “a” to complete “expanse” or “l” to form “explore”.

The longer the context window, the better the model can maintain coherence, but the more computer memory and processing time it requires.

Training Example

  • Input Sequence: “Hello my nam”

  • Target Sequence: “ello my name”

The model learns:

  • Given “H”, predict “e”
  • Given “He”, predict “l”
  • Given “Hel”, predict “l”
  • And so on…

This means our dataset of 5.8 million characters becomes about 47,000 chunks and, since every position in a chunk is a prediction, around 5.7 million individual training examples.

Creating Training Data

seq_len = 120

# A stream of character indices, cut into chunks of seq_len + 1
char_dataset = tf.data.Dataset.from_tensor_slices(encoded_text)
sequences = char_dataset.batch(seq_len + 1, drop_remainder=True)

def create_seq_targets(seq):   # "Hello my name"
    input_txt = seq[:-1]       # "Hello my nam"
    target_txt = seq[1:]       # "ello my name"
    return input_txt, target_txt

dataset = sequences.map(create_seq_targets)
dataset = dataset.shuffle(10000).batch(128, drop_remainder=True)

Character Embeddings

From Sparse to Dense Representation

Initially, each character is represented as a one-hot vector, which is mostly zeros with a single one indicating which character it is. For 123 characters, this means each character is represented by a vector with 123 elements, where 122 are zero and 1 is one. This is wasteful and doesn’t capture any relationships between characters.

Character embeddings solve this problem by representing each character as a dense vector of real numbers. Instead of 123 mostly-zero values, each character becomes 256 meaningful numbers. These numbers are learned during training and end up encoding relationships between characters.

Learning Character Relationships

Something fascinating happens during training: characters that behave similarly end up with similar embedding vectors. Vowels tend to cluster together because they can often substitute for each other in similar contexts. Consonants that frequently appear together (like ‘th’ or ‘ch’) develop related embeddings.

The model learns that uppercase and lowercase versions of the same letter are related but distinct. It discovers that digits form their own cluster since they appear in similar contexts (dates, measurements, chapter numbers). Punctuation marks develop embeddings based on their grammatical functions.

Visualization and Understanding

When we project these 256-dimensional embeddings down to 3D space for visualization, we can see these learned relationships. The embedding space becomes a map where distance represents similarity. Characters that can substitute for each other in many contexts end up close together, while characters with completely different roles end up far apart.

This learned representation becomes the foundation for everything else the model does. The quality of these embeddings directly affects the model’s ability to generate coherent text.

You can play with Word2Vec - Embedding Projector

Model Architecture

RNN Architecture Components

Our Jules Verne Bot consists of three main components, each serving a specific purpose in the text generation pipeline.

Embedding Layer: This is our translation layer. It takes character indices (numbers like 47, 83, 72) and converts them into dense 256-dimensional vectors that capture character relationships. Think of this as converting raw symbols into a format that captures meaning and relationships.

GRU Layer: This is the brain of our operation. With 1024 hidden units, this layer processes sequences and maintains memory about what it has seen. When processing the sequence “The submarine descended”, the GRU maintains a hidden state that encodes information about the submarine, the action of descending, and the overall maritime context.

Dense Output Layer: This is our decision-making layer. It takes the GRU’s 1024-dimensional hidden state and converts it into 123 probabilities, one for each character in our vocabulary. These probabilities represent the model’s confidence about what character should come next.

Model Summary

Model: "sequential_4"
_________________________________________________________________
Layer (type)                 Output Shape              Param #   
=================================================================
embedding_4 (Embedding)      (1, 120, 256)            31,488    
_________________________________________________________________
gru_3 (GRU)                  (1, 120, 1024)           3,938,304 
_________________________________________________________________
dense_3 (Dense)              (1, 120, 123)            126,075   
=================================================================
Total params: 4,095,867 (15.62 MB)
Trainable params: 4,095,867 (15.62 MB)
Non-trainable params: 0 (0.00 B)

Our model has 4,095,867 parameters. These are the individual numbers that the model adjusts during training to improve its predictions. To put this in perspective, each parameter is like a tiny dial that affects how the model processes information.

Training involves adjusting all 4 million dials to work together harmoniously.

The GRU layer contains most of these parameters (about 3.9 million) because it needs to learn complex patterns about how characters relate to each other across different time steps. The embedding layer has about 31,000 parameters (123 characters × 256 dimensions), and the output layer has about 126,000 parameters.

Memory and Processing Flow

When processing text, information flows through the model like this:

A character index enters the embedding layer and becomes a 256-dimensional vector. This vector enters the GRU, which combines it with its current memory state to produce a new 1024-dimensional hidden state. This hidden state captures everything the model “knows” at this point in the sequence.

The hidden state goes to the dense layer, which produces a score for each of the 123 possible next characters. Those scores become probabilities, and the next character is drawn from them. Always taking the most probable one is possible too (greedy decoding), but it quickly falls into loops; the Temperature Control section below shows how the draw is tuned.

Crucially, the GRU’s hidden state becomes its memory for the next character prediction. This creates a chain of memory that allows the model to maintain context across the entire sequence.

Why GRU over Basic RNN?

GRU Advantages:

  • Solves Vanishing Gradient: Better information flow through long sequences
  • Selective Memory: Can choose what to remember and forget
  • Computational Efficiency: Fewer parameters than LSTM
  • Better Performance: More stable training than basic RNNs

Training Process: Teaching the Model to Write

The Learning Objective

Training a neural network means adjusting its millions of parameters so it makes better predictions. We use a loss function called sparse categorical crossentropy, which measures how far off the model’s predictions are from the correct answers.

Think of it like teaching someone to play darts. Each throw (prediction) has a target (the correct next character). The loss function measures how far each dart lands from the bullseye. Training adjusts the player’s technique (the model’s parameters) to improve accuracy over time.

Hardware and Time Requirements

We trained our model on a Tesla T4 GPU, which can perform thousands of calculations simultaneously. This parallelization is crucial because each training step involves matrix multiplications with millions of numbers. The training took 33 minutes for 30 complete passes through the entire dataset.

To understand why we need a GPU, consider that training involves calculating gradients for all 4 million parameters, potentially thousands of times per second. A regular CPU would take many hours to complete the same training that a GPU accomplishes in minutes.

Monitoring Progress

During training, we watch the loss decrease from about 2.4 in the first epoch to about 0.87 by epoch 30. This represents the model’s improving ability to predict the next character. Early in training, the model makes essentially random predictions. By the end, it has learned sophisticated patterns about English spelling, grammar, and Jules Verne’s writing style.

The learning curve typically shows rapid improvement in the first few epochs as the model learns basic patterns like common letter combinations. Later epochs show slower but steady improvement as the model refines its understanding of more complex patterns like narrative structure and thematic elements.

Preventing Overfitting

One challenge in training is overfitting, in which the model memorizes the training data rather than learning generalizable patterns. The usual defense is to hold part of the text out, monitor the loss on it (the validation loss), and keep the model from the epoch where that loss was lowest.

The original training run did not do this: all ten books were used for training, so the loss we watched was measured on text the model had already seen. When the same model was later retrained with a held-out set (see From RNN to Transformer), the validation loss bottomed out at epoch 10 of 30 and rose from there, while the training loss kept falling. Two-thirds of the original training time was spent memorizing.

Training Configuration

Hardware Setup:

  • GPU: Tesla T4 (Google Colab)
  • GPU RAM: 15.0 GB
  • Training Time: 33 minutes for 30 epochs

Training Parameters:

  • Loss Function: Categorical Sparse Crossentropy
  • Optimizer: Adam (adaptive learning rate)
  • Epochs: 30
  • Batch Size: 128
  • Buffer Size: 10,000 (for dataset shuffling)

Training Implementation

The original notebook compiles and trains the model like this:

def sparse_cat_loss(y_true, y_pred):
    return sparse_categorical_crossentropy(y_true, y_pred, from_logits=True)

model.compile(optimizer=Adam(learning_rate=0.001), loss=sparse_cat_loss)
history = model.fit(dataset, epochs=30)

With a held-out set, two changes keep the best model instead of the last one: validation_data reports the loss on unseen text after every epoch, and ModelCheckpoint saves the model only when that loss improves. (Keras’s validation_split does not work here, because the data is a tf.data.Dataset.)

checkpoint = keras.callbacks.ModelCheckpoint(
    'rnn-split.keras', monitor='val_loss', save_best_only=True)

history = model.fit(train_ds, validation_data=val_ds,
                    epochs=30, callbacks=[checkpoint])

Text Generation

The Generation Process

Once trained, our model becomes a text generation engine. We start with a seed phrase like:

“THE FLYING SUBMARINE”

and ask the model to continue the story. The process works character by character:

The model receives “THE FLYING SUBMARINE” and predicts the most likely next character based on everything it learned from Jules Verne’s works. Maybe it predicts a space, starting a new word. Then we feed “THE FLYING SUBMARINE” (with the space) back to the model and ask for the next character.

This process continues indefinitely, with each new character becoming part of the context for predicting the next one. The model might generate “THE FLYING SUBMARINE descended into the mysterious depths…” as it draws upon patterns learned from Verne’s nautical adventures.

Temperature Control

Here’s where we can control the model’s creativity through a parameter called temperature. Temperature affects how the model chooses between different possible next characters.

With temperature set to 0.1, the model almost always picks the most probable next character. This produces very predictable, conservative text that closely mimics the training data but might be repetitive or boring.

With temperature set to 1.0, the model considers all possible next characters according to their learned probabilities. This produces more varied and creative text, but sometimes makes unusual choices that lead to interesting narrative directions.

With temperature above 1.5, the model becomes quite random, often producing text that starts coherently but gradually becomes nonsensical as unlikely character combinations accumulate.

In short:

  • Temperature = 0.5: More predictable, conservative text
  • Temperature = 1.0: More creative, diverse text
  • Temperature = 1.5: Very random, potentially nonsensical text

Implementation

def generate_text(model, start_string, num_generate=1000, temperature=1.0):
    # Convert start string to numbers
    input_eval = [char_to_idx[s] for s in start_string]
    input_eval = tf.expand_dims(input_eval, 0)

    text_generated = []
    for i in range(num_generate):
        predictions = model(input_eval)
        predictions = tf.squeeze(predictions, 0)

        # Apply temperature, then draw the next character
        predictions = predictions / temperature
        predicted_id = tf.random.categorical(predictions, num_samples=1)[-1, 0].numpy()

        # Append it to the input, so the model sees the whole text so far
        input_eval = tf.concat([input_eval, tf.expand_dims([predicted_id], 0)], axis=-1)
        text_generated.append(idx_to_char[predicted_id])

    return start_string + ''.join(text_generated)

The model is not stateful, so the input has to carry the text written so far: feeding it only the last character would leave it with a single character of context. Re-running the whole sequence for every new character is simple but gets slower as the text grows; the browser demo in the From RNN to Transformer section carries the GRU’s hidden state forward instead, at a constant cost per character.

Generation Example (Temperature = 0.5)

Seed: “THE FLYING SUBMARINE”

Generated Text:

THE FLYING SUBMARINE
CHAPTER 100 VENTANTILE

This eBook is for the use of anyone anywhere in the United States and most 
other parts of the earth and miserable eruptions. The solar rays should be 
entirely under the shock of the intensity of the sea. We were all sorts. Are 
we to prepare for our feelings?"

"I can never see them a good geographer," said Mary.

"Well, then, John, for I get to the Pampas, that we ought to obey the same
time. In the country of this latitude changed my brother, and the
_Nautilus_ floated in a sea which contained the rudder and
lower colour visibly. The loiter was a fatalint region the two
scientific discoverers. Several times turning toward the river, the cry
of doors and over an inclined plains of the Angara, with a threatening
water and disappeared in the midst of the solar rays.

The weather was spread and strewn with closed bottoms which soon appeared
that the unexpected sheets of wind was soon and linen, and the whole
seas were again landed on the subject of the natives, and the prisoners
were successively assuming the sides of this agreement for fifteen days
with a threatening voice.
...

Notice the second paragraph: “This eBook is for the use of anyone anywhere in the United States…” That is not Verne; it is the opening of the Project Gutenberg license, which appears in every book file. The model learned it like any other text.

Example Output Analysis

Let’s examine some generated text: “The weather was spread and strewn with closed bottoms which soon appeared that the unexpected sheets of wind was soon and linen, and the whole seas were again landed on the subject of the natives…”

This excerpt shows both the model’s strengths and limitations. It successfully captures Verne’s descriptive style and maritime vocabulary (“seas,” “wind,” “natives”). The sentence structure feels appropriately Victorian and elaborate. However, the meaning becomes confused with phrases like “closed bottoms” and “sheets of wind was soon and linen.”

This illustrates the fundamental challenge of character-level generation: the model learns local patterns (how words are spelled, common phrases) much better than global coherence (logical narrative flow, consistent meaning).

Challenges and Limitations

Context Window Constraints

Our 120-character context window creates a fundamental limitation. The model can only “see” about one paragraph of previous text when making predictions. This means it might introduce a character named Captain Smith, then 200 characters later introduce another character with the same name, having “forgotten” the first introduction.

Humans writing stories maintain mental models of characters, plot lines, and world-building details across entire novels. Our model’s memory effectively resets every 120 characters, making long-term narrative consistency nearly impossible.

Character vs. Word Level Trade-offs

Character-level generation requires many more prediction steps than word-level generation. Generating the phrase “extraordinary adventure” requires 22 character predictions instead of just 2 word predictions. This makes character-level generation much slower and more computationally expensive.

However, character-level generation offers unique advantages. The model can generate new words it has never seen before by combining character patterns. It can handle misspellings, made-up words, or technical terms more gracefully than word-level models that have fixed vocabularies.

Coherence Challenges

Perhaps the biggest limitation is maintaining semantic coherence. The model might generate grammatically correct text that makes no logical sense. It can describe “The submarine floating in the air above the mountain peaks” because it has learned that submarines float and that Verne often described mountains, but it hasn’t learned the physical constraint that submarines float in water, not air.

This happens because the model learns statistical patterns without understanding meaning. It knows that certain word combinations are common without understanding why they make sense.

Summary

  1. Limited Context Window
    • Issue: Only 120 characters of context
    • Impact: Cannot maintain coherence over long passages
    • Example: May forget characters or plot points mentioned earlier
  2. Character vs Word Level
    • Issue: Character-level generation is slower and less efficient
    • Impact: Requires more computation for equivalent output
    • Trade-off: Better handling of rare words vs efficiency
  3. Coherence Problems
    • Issue: May generate grammatically correct but semantically inconsistent text
    • Cause: Limited understanding of story structure and plot consistency
  4. Repetitive Patterns
    • Issue: May fall into repetitive loops
    • Cause: Model overfitting to common patterns in training data

Potential Improvements

  1. Longer Context Windows: Increase sequence length for better coherence
  2. Hierarchical Models: Separate models for different text levels (word, sentence, paragraph)
  3. Fine-tuning: Additional training on specific styles or topics
  4. Beam Search: Better text generation algorithms instead of greedy sampling

These limitations are what led to modern Language Models based on the Transformer architecture. Before looking at them at scale, the next section asks a narrower question: with everything else held equal, what does a Transformer change?

From RNN to Transformer: A Fair Comparison

The previous sections end with a claim you will find in almost every introduction to language models: Transformers replaced RNNs because attention handles context better. It is true at the scale of GPT-3. But does it hold for our small model, on our ten books? This section answers that with an experiment you can reproduce, and the answer is more interesting than “the Transformer wins.”

Changing One Thing at a Time

Three things separate the Jules Verne Bot from a modern language model: architecture (recurrence against attention), scale (4 million parameters against billions), and tokenization (characters against subwords). Change all three at once, and a better result tells you nothing: nobody could say which change did the work.

So the experiment holds two of them fixed and changes one at a time:

  • Same tokenizer: the same 123 characters, built the same way from the same books.
  • Same size: the Transformer is built to have almost exactly as many parameters as the RNN — 4.01 million against 4.10 million, 2% apart.
  • Same text, same test: both are trained on the same books and measured on the same held-out pages.
Model Parameters Context window What it tests
RNN (GRU) 4,095,867 120 the baseline
Transformer 4,011,520 120 architecture only
Transformer 4,046,336 256 architecture + a longer window

The first two rows are a fair fight: same size, same window, same data. The only difference left is how each model carries the past.

Two Ways to Remember

The RNN compresses everything it has read into one fixed-size vector, the hidden state, and updates it at every character. That is its memory, and it is also its limit: whatever does not fit in 1,024 numbers fades.

The RNN used in this chapter: embedding, a 1,024-unit GRU, a dense layer, and sampling. The hidden state carried from one character to the next is the model’s memory.

The Transformer keeps nothing compressed. For every character in its window, it stores a key and a value, and each new character looks back at all of them at once and decides which ones matter — that is attention. Our Transformer is a small, decoder-only model: five blocks, four attention heads, and 256-dimensional vectors, sized to match the RNN.

The size-matched Transformer. Each character attends to every earlier one in the window; the KV cache stores those keys and values so that generating a new character costs one pass instead of re-reading the whole window.

The two architectures side by side: how each carries the past, what each character costs to generate, and what was measured.

Preparing a Fair Test

Two changes to the data were needed before the comparison could mean anything.

Removing the Gutenberg boilerplate. Every book file begins with a Project Gutenberg header and ends with the full license — 3.4% of the text. That is how the license ended up in the RNN’s generated sample earlier in this chapter. A cleaning script keeps only the text between the *** START OF THE PROJECT GUTENBERG EBOOK *** and *** END *** markers, leaving 5,571,763 characters. The vocabulary is still built from the original files, so it stays at 123 characters, and every model remains compatible.

Holding out text neither model has seen. From the middle of each of the ten books, a continuous 5% slice is set aside: 278,888 characters in total. No model trains on it. The loss measured on this text, the validation loss, tells us how well a model predicts Verne it has never read, rather than how well it remembers the pages it trained on. The original RNN had no such test; it was retrained here on exactly the same split.

Every model also keeps its best checkpoint, the one with the lowest validation loss, rather than the last one. As we saw earlier, for the RNN that was epoch 10 of 30.

Reading the Results

Here is what the three models scored on the held-out text:

Model Window Validation loss Perplexity
RNN (GRU) 120 1.146 3.15
Transformer 120 1.102 3.01
Transformer 256 1.034 2.81

Before interpreting them, it is worth knowing what these numbers mean.

Loss. At every position, the model assigns a probability to each of the 123 possible characters. The loss is the average of −ln(p), where p is the probability the model gave to the character that actually came next. A model that was always certain and always right would score 0. The unit is nats, because it uses the natural logarithm; dividing by ln 2 ≈ 0.693 converts it to bits. The RNN needs 1.65 bits to predict each character, the Transformer with the longer window 1.49.

Probability of the right character. Undoing the logarithm, e^(−loss) gives the geometric mean of the probability assigned to the correct character: about 0.32 for the RNN, 0.33 for the Transformer at 120, and 0.36 at 256.

Perplexity. It is e^loss, and it has an intuitive reading: the model is as uncertain as if it were choosing uniformly among that many characters. A model that knew nothing would pick among all 123 — perplexity 123, loss ln 123 ≈ 4.81, which is exactly where training starts (4.76 at step 0). All three models bring that down to about three candidates per character.

With that in mind:

  1. At the same size and window, attention wins, but only by a little. The loss drops from 1.146 to 1.102: 0.044 nats per character, about 4%. The difference is real — each model was trained with three different random seeds, and the seed-to-seed variation was only 0.002, so the gap is about 22 times the noise. But it is too small to see by reading a paragraph of each.
  2. A longer window helps more than the architecture change. Keeping the Transformer the same size and doubling its window to 256 characters lowers the loss by another 0.068, more than attention itself bought.
  3. The RNN memorizes. Its validation loss was lowest at epoch 10 and rose afterward, while its training loss kept falling — a textbook case of overfitting that the original run, with no validation set, could not show.

So why did the Transformer replace the RNN, if it barely wins here? Because the advantage of attention is not mainly quality at a fixed small size. It is scale. An RNN must process text one character after another, even during training; a Transformer processes a whole window in parallel. That is what made it practical to train models with billions of parameters on trillions of tokens — and at that scale, the gap becomes enormous. At 4 million parameters and ten books, both models hit the same wall: the size of the corpus.

A Lesson from the KV Cache

Generating text with a Transformer uses a KV cache: the keys and values of the characters already written are stored, so each new character costs one pass through the model instead of re-reading the whole window. It is the Transformer’s version of the RNN’s hidden state.

The first version of this project got it wrong in an instructive way. When the text grew past the window, the cache was used as a ring, with each new character overwriting the oldest. That seems natural, but our Transformer uses learned absolute positions: each stored key carries the position it was written at. Once the ring wrapped, the newest character sat in the oldest slot, with the lowest position number, and the model read a rotated version of the text. Past the window, the output collapsed into word salad.

The fix is to rebuild the cache when the window fills: drop the oldest 64 characters and recompute the rest at their new positions. Measured by the share of real words written past the window, the text went from 54% to 97%. The lesson generalizes: a check that only compares outputs within the first window could not catch this, and neither could a test that only checked the length of the output. The test now counts words that exist in the corpus.

Modern models avoid the problem differently, with relative position schemes such as RoPE and ALiBi, which make a sliding window valid by construction.

Try It in the Browser

All three models run in a web page, with no server and no framework: the weights are downloaded once (about 8 MB each), and the forward passes are written in plain JavaScript. Run all three writes from the same seed with each model, so the numbers above can be read next to the text they correspond to.

Going Further: Experiments for You

This experiment answers one question and leaves several open. Each of these can be run on a free Colab GPU with the notebooks in the repository:

  1. Scale. Retrain the 25-million-parameter Transformer (--preset large) on the cleaned corpus and the shared split. Does a model six times larger finally separate from the RNN, or does the small corpus hold it back? Watch for the epoch where its validation loss stops improving.
  2. Give the RNN the same chance. The RNN was trained on 120-character sequences. Train it on 256-character sequences and compare it with the 256-character Transformer. Is the gain from the longer window, or from attention using it?
  3. Close the regularization gap. The Transformer uses dropout (0.1) and the RNN uses none. Add dropout to the GRU and retrain. How much of the 0.044 remains?
  4. Repeat the 256-character run. Only the 120-character models were run with three seeds. Run two more seeds of the 256-character Transformer to confirm the 0.068.
  5. Change the tokenizer. Keep the architecture and the size, and switch from characters to subwords (for example, a small BPE vocabulary). This is the third variable of the three, and the one this chapter did not touch.
  6. Change the author. Train the same pair on another corpus — Machado de Assis in Portuguese, for example — and see whether the conclusions hold in another language.

Whichever you try, keep the rule that made this comparison meaningful: change one thing at a time, and measure on text the model has never seen.

The Transformer comparison, the browser demo, and the training and export scripts were generated by DeepSeek 4.1 Flash and reviewed by Claude Opus 5 (Anthropic) under the author’s direction, in September 2026. The original RNN notebook dates from 2024.

Connecting to Modern Language Models

Scale Comparison

To appreciate how far language modeling has advanced, consider the scale differences between our RNN Jules Verne Bot and modern language models:

Our model has 4 million parameters and was trained on about 5.8 million characters (10 books). GPT-3 has 175 billion parameters and was trained on about 570 GB of filtered text — some 300 billion tokens, selected from 45 terabytes of raw web crawl. That is over 40,000 times more parameters and about 100,000 times more training text.

Modern small language models (SLMs) like Phi-3-mini still dwarf our model with 3.8 billion parameters, but they represent more efficient designs that achieve impressive performance with “only” 1,000 times more parameters than our model.

Aspect Jules Verne Bot GPT-3 (2020) Phi-3-mini (2024)
Architecture RNN (GRU) Transformer Transformer
Parameters 4 million 175 billion 3.8 billion
Training Data 5.8M characters ~570 GB (300B tokens) 3.3T tokens
Context Length 120 characters 2,048 tokens 128,000 tokens
Tokenization Character-level (123) Subword BPE (50,257) Subword (32,064)
Training Time 33 minutes Weeks 7 days
GPU Requirements 1 Tesla T4 Thousands of V100 GPUs 512 H100 GPUs

Architectural Evolution

The biggest advancement since RNNs is the Transformer architecture, which uses attention mechanisms instead of recurrent processing. While RNNs process text sequentially (like reading word by word), Transformers can examine all parts of a text simultaneously and learn relationships between any two words, regardless of how far apart they are.

This addresses the long-term memory problem that limits our RNN model — within the window. A Transformer with a 128,000-token context can keep a character introduced in “chapter 1” in view while writing “chapter 10”, something our 120-character window makes impossible. But the window still ends somewhere, and anything beyond it is as invisible to a Transformer as it is to our RNN.

Training Efficiency

Modern models also benefit from more sophisticated training techniques. They’re pre-trained on massive, diverse datasets to learn general language patterns, then fine-tuned on specific tasks. They use techniques like instruction tuning, where they learn to follow human commands, and reinforcement learning from human feedback, where they learn to generate text that humans find helpful and appropriate.

Summary: Why Modern Models Perform Better?

  1. Transformer Architecture
    • Attention Mechanism: Can look at any part of the input sequence
    • Parallel Processing: Much faster training and inference
    • Better Long-range Dependencies: Maintains context over thousands of tokens
  2. Scale
    • More Data: Trained on vastly more diverse text
    • More Parameters: Can memorize and generalize better
    • More Compute: Allows for more sophisticated training techniques
  3. Advanced Techniques
    • Pre-training + Fine-tuning: Learn general language then specialize
    • Instruction Tuning: Trained to follow human instructions
    • RLHF: Reinforcement Learning from Human Feedback

Conclusion

Building the Jules Verne Bot teaches us that creating artificial intelligence systems capable of generating human-like text requires careful consideration of multiple components working together. The embedding layer learns to represent characters meaningfully, the RNN layer processes sequences and maintains memory, and the output layer makes predictions based on learned patterns.

The project also illustrates the fundamental trade-offs in machine learning: between model complexity and training speed, between creativity and coherence, between local accuracy and global consistency. These trade-offs appear in every AI system, from simple character-level generators to the most sophisticated language models.

Most importantly, this project demonstrates that impressive AI capabilities emerge from relatively simple components combined thoughtfully. Our 4-million parameter model, while limited compared to modern systems, genuinely learns to write in Jules Verne’s style through nothing more than statistical pattern recognition and mathematical optimization.

The techniques we’ve explored, sequence processing, embedding learning, and generation strategies, form the foundation for understanding any language model. Whether you encounter RNNs, Transformers, or future architectures yet to be invented, the core concepts remain consistent: learn patterns from data, encode meaning in mathematical representations, and generate new content by predicting what should come next.

Understanding these fundamentals provides the foundation for working with, improving, or creating the next generation of language models that will shape how humans and computers communicate in the future.

Resources