RRishiinrishi2220.hashnode.dev·Jul 22 · 36 min readTowards the Modern Transformer ArchitectureHi y'all, Long time no see! Last post was last year. Everyone has likely read Attention Is All You Need, and understood everything around self-attention, and mapped out how encoders and decoders talk 11K
JJessenindescendingnotebooks.com·Apr 1, 2025 · 16 min readPositional Encoding from Sinusoidal to RoPETransformers process the tokens of a text input in parallel, but unlike sequential models they do not understand position and see the input as a set of tokens. However when we calculate attention for a sentence, words that are the same but in differe...00