Knowledge Graph — Coursera NotesAcademic disciplinesInformation Technology / Computer ScienceArtificial IntelligenceDeep Learning

Transformer

concept · part of Deep Learning

Underlying architecture of modern LLMs

They use self-attention mechanisms to capture intricate relationships within data sequences, enabling human-like language generation. Examples include GPT for text generation.

Transformers process all tokens in parallel rather than sequentially, which enables efficient training on large datasets. They consist of an encoder and decoder stack, with each layer containing multi-head self-attention and feed-forward neural networks. Positional encodings are added to input embeddings to retain sequence order information.

Inside Transformer (2)

Connections

This is the text view of an interactive 3D knowledge graph — open this page with JavaScript enabled to explore it visually.

🧠 Knowledge Graph

Select a node

The owner's editing tools — shown here so you can see how the graph is grown, but read-only.

Click a bubble to drill in · click again to collapse · drag to move around