20 points | by cdnsteve 2 hours ago ago
5 comments
Terminal Bench 2/2.1 appears to be almost solved, so we probably shouldn't look too hard on that? Their Terminal Bench 4 score is soso.
Still quite impressive though.
These tests are pointless when often models get nerfed few days after release. Astra today is way dumber than just few days ago.
Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?
Not Astra but Sol hasn't been nerfed: https://marginlab.ai/trackers/codex/
How and why do they get nerfed? To save money?
Terminal Bench 2/2.1 appears to be almost solved, so we probably shouldn't look too hard on that? Their Terminal Bench 4 score is soso.
Still quite impressive though.
These tests are pointless when often models get nerfed few days after release. Astra today is way dumber than just few days ago.
Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?
Not Astra but Sol hasn't been nerfed: https://marginlab.ai/trackers/codex/
How and why do they get nerfed? To save money?