Knowledge Graph — Coursera NotesAcademic disciplinesInformation Technology / Computer ScienceArtificial IntelligenceLarge Language ModelsFoundation Models

RoBERTa

concept · part of Foundation Models

Facebook variant of BERT trained with more data and compute, dropping next-sentence prediction.

RoBERTa uses dynamic masking, larger mini-batches, and a larger byte-level BPE vocabulary, and it is trained on 160GB of text (including CC-News, OpenWebText, and Books) compared to BERT's 16GB.

Connections

This is the text view of an interactive 3D knowledge graph — open this page with JavaScript enabled to explore it visually.

🧠 Knowledge Graph

Select a node

The owner's editing tools — shown here so you can see how the graph is grown, but read-only.

Click a bubble to drill in · click again to collapse · drag to move around