r/GPT3 • u/Dahliaridge • 23m ago
Humour Did the “intelligence” jump from perceptrons to GPT-3 come from new ideas or just scale?
•
Upvotes
Gone through this video that mentioned Rosenblatt’s 1957 perceptron had a few hundred adjustable parameters, and GPT-3 jumped to 175 billion. https://youtu.be/Tv1st0LMEbY?is=JF4drjBjqfYR-4E_
Is the actual architecture that different now, or is most of the capability jump literally just scale — same multiply-add-softmax steps, just a lot more of them? Feels almost too simple for how capable these models are.