28 points | by stared 2 hours ago ago
2 comments
“We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.”
Uh.. okay.. but whats a run… read blog
“We want to measure how well frontier models can conduct research….””we ran 153 autonomous runs on the nanoGPT optimizer speedrun across”
Okay but what is a optimiser run and what connection does it have to being good at research?
“For comparison, Anthropic's internal automated AI R&D evaluation optimizes a model on a CPU node,”
So I should go look what Anthropic was doing to understand?
Why not just explain what it means in their blog..
Neat!
The graphs show the "best validated result" for each model. I wonder how much variation there is between runs for a model?
“We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.”
Uh.. okay.. but whats a run… read blog
“We want to measure how well frontier models can conduct research….””we ran 153 autonomous runs on the nanoGPT optimizer speedrun across”
Okay but what is a optimiser run and what connection does it have to being good at research?
“For comparison, Anthropic's internal automated AI R&D evaluation optimizes a model on a CPU node,”
So I should go look what Anthropic was doing to understand?
Why not just explain what it means in their blog..
Neat!
The graphs show the "best validated result" for each model. I wonder how much variation there is between runs for a model?