Facebook variant of BERT trained with more data and compute, dropping next-sentence prediction.
RoBERTa uses dynamic masking, larger mini-batches, and a larger byte-level BPE vocabulary, and it is trained on 160GB of text (including CC-News, OpenWebText, and Books) compared to BERT's 16GB.
This is the text view of an interactive 3D knowledge graph — open this page with JavaScript enabled to explore it visually.
🧠 Knowledge Graph
Select a node
The owner's editing tools — shown here so you can see how the graph is grown, but read-only.
Click a bubble to drill in · click again to collapse · drag to move around
Quiz
Proposed changes
⚠️ Includes destructive changes (merge/delete). They can still be reverted with Undo after applying.
🔒 Only the owner can edit this graph
You're browsing a read-only snapshot. The editing tools are shown so you can see how the graph is grown, but changes are disabled here — exploring, searching and path-finding all work.