In early 2017, software engineers inside Google’s headquarters were stuck on a persistent efficiency problem. Google Translate relied on Recurrent Neural Networks, systems that processed text sequentially, one word at a time. To translate a sentence from German to English, the model read the first word, calculated its internal representation, and then moved to the second word.
This strict step-by-step sequence created a major computational bottleneck. Graphics processing units, built to execute thousands of mathematical calculations simultaneously, sat underutilized while waiting for sequential loops to finish. Longer sentences made the problem worse.
By the time a model reached the end of a long paragraph, context from the opening sentences often degraded. Engineers tried using Long Short-Term Memory networks to preserve earlier context, but training times remained long and translation quality dropped on complex documents.
Google’s 2017 “Attention Is All You Need” Paper Quietly Started a New AI Age: Here’s How It Changed Everyday Life
The breakthrough came from abandoning sequential processing altogether. Instead of reading text left to right, the new architecture—named the Transformer—fed entire blocks of text into memory simultaneously. It tracked connections between words using a mathematical calculation known as self-attention.
Self-attention assigns numerical weights to every word pair in a passage. In the sentence "The animal didn't cross the street because it was too tired," the system calculates matrix probabilities to determine whether "it" refers to the animal or the street. It does this by creating three numerical vectors for every word token: a query, a key, and a value. Multiplying the query matrix by the key matrix generates an attention score, measuring how much weight every word gives to every other word across the sentence at once.
Because matrix multiplication runs efficiently across thousands of GPU cores, training speeds accelerated. Models could process vastly larger training datasets without stalling on word-by-word feedback loops.
Eight Authors and an Open Source Publication
In June 2017, eight researchers published their design on the arXiv preprint server under the title Attention Is All You Need . The co-authors—Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin—had originally designed the framework to improve performance on standard English-to-German and English-to-French translation benchmarks.
Rather than keeping the architecture proprietary, Google made the paper and code publicly available. That decision reshaped the technology sector. Within six years, all eight co-authors had departed Google to launch their own prominent startups, including Character.AI, Cohere, Inceptive, Essential AI, and Sakana AI. The eight-page document became one of the most cited research papers in modern computer science, providing a universal blueprint for industrial research labs worldwide.
From Grammar to General Computation
Researchers quickly realized the Transformer architecture was not limited to language translation. In 2018, Google introduced BERT, using bidirectional Transformer layers to process search queries. Around the same time, OpenAI adapted the decoder mechanism to build GPT-1, focusing on predicting the next word in a sequence.
The underlying math remained identical to the 2017 blueprint, but scaling parameter counts changed what the models could do. Increasing model sizes from hundreds of millions of parameters to hundreds of billions revealed an unexpected property: training a Transformer to predict the next token across massive text datasets produced emergent abilities in logic, software coding, and textual synthesis. The same matrix calculations originally designed to align vocabulary between two languages were now writing functional Python code, analyzing complex legal filings, and predicting protein structures in molecular biology.
The widespread adoption of Transformers transformed global hardware manufacturing and energy demand. Because self-attention requires high memory bandwidth to manage large context windows, chip designers modified their production pipelines. Nvidia added dedicated tensor cores to its hardware specifically to accelerate the matrix operations used in Transformer layers.
Training modern foundation models now requires tens of thousands of specialized chips clustered inside specialized data centers. Memory bandwidth and inter-chip connectivity have replaced clock speed as the primary metrics for hardware performance.
Power consumption for these compute clusters has risen sharply, forcing technology companies to secure dedicated nuclear and solar power purchase agreements. An engineering fix originally written to clean up translation latency ended up driving global semiconductor design and industrial energy investment.