Skip to main content

One doc tagged with "gpt"

View all tags

Transformer Variants

One architecture, three ways to cut it — and which cut you choose determines what the resulting model can actually do. The difference between a model that reads and a model that writes turns out to come down to one thing: the shape of the attention mask.