How transformers changed AI forever
The Transformer did more than improve language models. It made it possible to train AI systems that understand context at extraordinary scale, laying the foundations for today’s large language models and generative AI.
For years, artificial intelligence systems struggled with a basic problem: how do you make sense of context? Words in a sentence, events in a customer journey, transactions in a financial system, pixels in an image, proteins in a cell.
The meaning of any one thing is often shaped by what came before, what sits alongside it, and what might follow. For most of that time, the assumed answer was sequence. Language runs in order, so the models processed it in order too. That assumption held until 2017, when a research paper with a deceptively plain title, Attention Is All You Need, introduced the Transformer.
It changed the technical trajectory of AI, and eventually the everyday experience of millions of people using generative AI. It is now among the most cited works in the field, referenced in more than a quarter of a million academic papers on Google Scholar. I think it is still worth reading, and not only for the historical significance.
Before transformers
The neural language models that immediately preceded the Transformer, recurrent neural networks and their more capable relatives, long short-term memory networks, processed text one word at a time and had to work through sequences in order.That made intuitive sense. Language is sequential, after all. But it came with serious trade-offs:
- Long-range relationships were difficult to retain. If the important clue appeared 500 words ago, the model could simply lose track of it.
- Training was slow. Processing could not be easily parallelised, so the hardware sat underused.
- Context degraded. The further back the relevant words sat, the weaker the signal became.
- Scaling was awkward. More data and more compute bought less than they should have.
The breakthrough: attention
Transformers took a different approach. Rather than treating every word as something that must be processed strictly in sequence, they allowed the model to consider relationships across the whole input. This mechanism is called self-attention.
Attention itself was not new. Bahdanau and colleagues had used it to improve machine translation in 2014. The contribution of the Transformer was to dispense with recurrence and convolution entirely and rely on attention alone, which is what made the architecture so well suited to modern computing hardware.
As the paper reports, the base model trained in twelve hours on eight graphics processing units, and the largest model in three and a half days, reaching state-of-the-art translation results at a small fraction of the training cost of comparable systems. In plain terms, self-attention helps a model weigh which other parts of a sequence matter when interpreting a particular token.
Take the much cited illustration from The Illustrated Transformer, "The animal did not cross the street because it was too tired". Attention heads in trained models can learn to place weight on the animal when processing "it", and that is part of what makes coreference tractable at scale. It is a useful illustration rather than a guarantee. Attention weights are not a reliable explanation of a model's output, and harder coreference cases still defeat current systems.Even so, attention gave models a far more practical way to handle context, and it made that context much cheaper to learn from than sequential processing allowed.
One qualification matters here. Attention cost grows quadratically with sequence length. In standard implementations, double the context and you quadruple the attention compute, which is why context length remained the central engineering constraint long after 2017, and why subquadratic and non-attention architectures, including state space models such as Mamba, remain an active field of research.The Transformer made context tractable at scale. It did not close the problem.
Why scale suddenly mattered
The Transformer did not create large language models by itself. But it supplied the architecture that made them feasible.Once researchers could train models on vast datasets using large clusters of graphics processing units, a pattern became impossible to ignore: bigger models, more data and more compute often produced substantially more capable systems.
Pre-training scale was never the whole story, but for a few years it looked like the main one. Those relationships have since been revised. Work published in 2022 showed that for compute-optimal training, model size and the number of training tokens should be scaled equally, so that doubling the model size calls for doubling the tokens. The claim that scaling produces sharp, unpredictable new abilities has also been contested on methodological grounds, on the basis that those apparent jumps tend to disappear under different metrics or better statistics.
Since 2024, much of the marginal investment has moved from pre-training scale into post-training and reasoning compute. This created the conditions for systems such as BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), Claude, Gemini and many others.
The question was no longer only, "Can we build a model that performs a task?" It became, "What capabilities emerge when we scale a general-purpose model?" That shift has influenced research priorities, investment flows, data centre construction, semiconductor supply chains, public policy, education and workplace technology.
Beyond language
Although Transformers first became famous for language, they quickly spread across AI. They are now used in areas including:
- Image generation and computer vision
- Speech recognition and synthesis
- Software development
- Scientific research and protein modelling, where AlphaFold produced structures that had previously taken laboratories years to determine
- Robotics and autonomous systems, including the Robotics Transformer family
- Cybersecurity analysis
- Financial modelling and anomaly detection
The deeper lesson is that the architecture was not narrowly about language. It was about finding patterns in complex, high-dimensional data where relationships matter, and that is a remarkably broad category.
The governance problem
Transformers also changed the governance challenge. Traditional AI systems were often designed for a discrete purpose: detect fraud, classify an image, recommend a product, forecast demand. Risks could be substantial, but the boundaries of the system were reasonably clear.
Foundation models are different. They are general-purpose systems trained on broad data at scale, and they can be adapted to an enormous range of uses, including uses their developers may never have anticipated.
The governance questions I hear most often are still the wrong ones. They ask whether a model is safe in the abstract, when the practical questions are about authority, evidence and accountability in particular settings.
- Who is accountable when a general-purpose model is deployed in a high-risk setting?
- What evidence is enough to demonstrate safety, reliability and suitability?
- How should organisations manage confidential data when employees use public AI tools?
- What does provenance mean when training data is collected at internet scale
- How can governments preserve capability and sovereignty when foundational infrastructure is controlled elsewhere?
These are not abstract questions. They are now procurement, risk, legal, security and capability questions for every serious organisation.
Nor are they entirely unaddressed. The European Union's AI Act imposes obligations on providers of general-purpose models. The NIST AI Risk Management Framework and ISO/IEC 42001 give organisations structures for governing AI systems.
Australia's path has been less linear than either, and more instructive. After consulting on mandatory guardrails for high-risk settings in 2024, the government declined to proceed with them, and the National AI Plan of December 2025 leaned instead on existing laws, sector regulators and voluntary guidance, supported by the Voluntary AI Safety Standard and the new AI Safety Institute.
By July 2026 the position had shifted again. In a speech at the University of Sydney, the Prime Minister announced mandatory Australian Standards for AI, drawing data centre expectations into a single regulatory framework, a new Office of AI inside the Department of the Prime Minister and Cabinet, and legislation the government aims to introduce in early 2027, presented as the first national framework of its kind. The same speech committed to copyright protections for Australian writers, artists and journalists, and to binding obligations on large data centres covering siting, grid connection, energy and water.
That reversal inside two years is worth reading closely. It suggests that the difficult question is no longer whether to regulate general-purpose systems, but how to do so without forfeiting the investment, the energy and the infrastructure those systems now depend on. What remains unsettled is not whether governance exists, but whether it is proportionate and enforceable.
AI is now infrastructure
It is tempting to describe generative AI as another software category. That misses entirely what has changed. The Transformer helped to turn AI into infrastructure. It sits behind tools for writing, coding, searching, translation, analysis, image creation and increasingly, agentic workflows.
Its effects will be felt not only through chat interfaces, but through the systems that shape decisions, services and the flow of information. For Australia, that means thinking beyond uptake. We need capability, assurance, secure data practices, meaningful evaluation and a clear view of where our dependencies sit. The Transformer made a new class of general-purpose model possible. That is the achievement, and it is a genuine one.
Whether the systems built on it serve the public interest is a separate question, and it will not be answered by scaling faster than our capacity to govern.
The architecture settled what was possible - it did not settle who is accountable for it.