- Start
- Jun 12, 201790% CONFIDENCEfrom the source
Google Publishes Attention Is All You Need
- "Attention Is All You Need" was posted to arXiv on 12 June 2017 by eight scientists and engineers at Google: Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser and Polosukhin[1][2]
- Its abstract begins: "The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration." The paper's answer was attention alone[1]
- The original encoder-decoder transformer had 100M parameters; the focus at the time was improving Seq2seq techniques, building on the attention mechanism proposed by Bahdanau et al. in 2014[2]
- The architecture "has become the main architecture of a wide variety of artificial intelligence systems, including large language models"[2]
- After publication, each of the eight authors left Google to join other companies or found startups[2]
References 388% CONFIDENCE
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%arxiv.org/abs/1706.03762arxiv.org· Posted Sep 22, 2026· Starts Jun 12, 2017 ✓· 33% of score
arXiv's own metadata for 1706.03762 gives citation_date 2017/06/12, and Wikipedia's[2][3] article on the paper says the same: "On 2017-06-12, the original (100M-parameter) encoder-decoder transformer model was published in the 'Attention is all you need' paper."
- [2]90%Attention Is All You Needen.wikipedia.org· Added Sep 22, 2026· 33% of score
Dates the paper to 12 June 2017 and describes what it did: "a 2017 research paper on machine learning authored by eight scientists and engineers working at Google" introducing the transformer, "based on the attention mechanism proposed in 2014 by Bahdanau et al.", which "has become the main architecture of a wide variety of artificial intelligence systems, including large language models". It also records that all eight authors later left Google.[1]
- [3]85%Transformer (deep learning architecture)en.wikipedia.org· Added Sep 22, 2026· 33% of score
The architecture itself, for how self-attention replaced recurrence and what was built on it afterwards.[1]
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most.