Extrapolating GPT-N performance
📖 Description
[Brown et al. (2020)](https://arxiv.org/pdf/2005.14165.pdf) (which describes the development of GPT-3) contains measurements of how 8 transformers of different sizes perform on several different benchmarks. In this post, I project how performance could improve for larger models, and give an overview of issues that may appear when scaling-up. Note that these benchmarks are for 'downstream tasks' that are different from the training task (which is to predict the next token); these extrapolations thus cannot be directly read off the scaling laws in OpenAI's Scaling Laws for Neural Language Models ([Kaplan et al., 2020](https://arxiv.org/pdf/2001.08361.pdf)) or Scaling Laws for Autoregressive Generative Modelling ([Henighan et al., 2020](https://arxiv.org/pdf/2010.14701.pdf)).
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +10 | Always |
| Vibey Doom | +2 | Always |
💭 Reactions
"Useful research for the community"
"Discussed in AI safety community"
🔗 Sources
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)