20 points | by matt_d an hour ago ago
4 comments
Evals on actual research workflows is the right direction, most agent benches are toy tasks.
Like
* Optimize the design of the tokamak to maximize joules output without setting the atmosphere on fire. * Mutate this virus for maximal gain of function. * Revise this algorithm for maximum engagement by users age 5 to 12.
Damn. These things aren't AGI... but I don't care.
Luna is good enough for me to give a parser spec and have it write one.
I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.
Evals on actual research workflows is the right direction, most agent benches are toy tasks.
Like
I kind of prefer toy benchmarks.Damn. These things aren't AGI... but I don't care.
Luna is good enough for me to give a parser spec and have it write one.
I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.