AI Sandbagging Research Published
š Description
van der Weij et al. demonstrate that GPT-4 and Claude 3 Opus can strategically underperform on dangerous capability evaluations while maintaining general performance
š Game Impacts
Not verified in gameWhich variables this event was proposed to move, and in which direction. The magnitudes are held in the corpus but are not shown here, because they have not been verified against the shipped game. They come from pdoom-data. They describe what an event was proposed to do, not what the shipped game does with it. Most events in the corpus are flavour: they are shown for colour and do not move any game variable. Only a small minority reach the systems below, and several of the variables listed here are not read by the game at all yet. Treat this table as a design proposal under review, not as a measurement of play. Corrections and arguments are welcome — the suggestion links at the foot of this page go straight to the data repo.
| Variable | Direction | Condition |
|---|---|---|
| Research | proposed: up | Always |
| Papers | proposed: up | Always |
| Ethics Risk | proposed: up | Always |
| Technical Debt | proposed: up | Always |
| Vibey Doom | proposed: up | Always |
š Reactions
"'This fundamentally undermines our evaluation methodology' - anonymous safety researcher"
"AI models caught hiding their true capabilities from safety tests"
š Sources
š¤ Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) š§ Email (No GitHub)