Turning point Draft

"Attention Is All You Need": the Transformer arrives, the idea ChatGPT will grow from

In June 2017, eight researchers, mostly from Google, proposed a new way to build neural networks: the Transformer. At the time it looked like a paper about machine translation. Today nearly every chatbot, ChatGPT included, is built on the idea.

To see why this paper matters, recall how computers handled text before. Neural networks read a sentence strictly in order, word by word, like a person tracing a line with a finger. It was slow, and by the end of a long text the model had trouble remembering the beginning.

On 12 June 2017, eight researchers, Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin, posted the paper "Attention Is All You Need." They proposed a neural network design they called the Transformer.

The Transformer looks at the whole text at once and, for each word, decides which other words to pay "attention" to in order to understand the meaning. In the sentence "the cat didn't jump onto the wardrobe because it was too high," for example, the model can work out that "it" is the wardrobe, not the cat.

The authors tested the idea on machine translation and beat the best systems of the day. Training took only a few days.

Why it matters

The new approach had one key advantage: the work could be split across many processors at once. That meant a model could be trained on huge amounts of text and made ever bigger, and it got smarter as it grew. Five years later, that recipe of "more data and more compute" led to ChatGPT.

The T in GPT stands for Transformer. But in 2017 the paper looked like a strong piece of work on translation; its real significance became clear only later.

For the curious

On the WMT 2014 (Workshop on Machine Translation, an annual machine translation competition) English-to-German benchmark, the model scored 28.4 BLEU (Bilingual Evaluation Understudy, an automatic measure of translation quality), more than 2 points above the previous record. On English-to-French it scored 41.8 BLEU, a single-model record, after 3.5 days of training on eight GPUs (graphics processing units). The Transformer dropped the recurrent and convolutional networks that had been the standard for working with text.

Sources

  1. Attention Is All You Need, arXiv, 12 June 2017 (primary source)