Just by seeing the title I knew it was going to be like the meme of Obama giving himself a medal.
Why are they even doing this? Are there still people that trust LLM benchmarks made by the LLM companies? It seems to me that the only people still using Claude are the ones that have it for free at work (me). Some of my colleagues even started using their own private OpenAI subscriptions to avoid using Claude, others are using Gemini flash to decipher what Claude is saying...
Yep, conflict of interest is the elephant in the room and its absence from the "reducing risks from advanced AI" point list is conspicuous:
* Figuring out how to prevent the incentives of frontier AI labs from aligning with anti-social deployment of AI rather than pro-social deployment of AI
It's not like the AI can simply advise them how to fix this because the labs already understood this risk perfectly well before they had an incentive not to. They put in place organizational structures to control it and then promptly smashed the structures once they smelled money. They already failed the integrity check. Even if their AI told them what they didn't want to hear I'm sure they would ignore it. Maybe they already have.
The underlying idea is that AI capabilities will become so advanced that only AI will enable us to monitor/correct/understand behavior. Obviously this is not without issue and I don't want to try to defend their position right now. But that's what they mean
That first sentence of yours explains exactly why it is so ridiculous. If only AI can understand it, how is there any assurance that AI will "correct" it's behavior that is aligned with what humans presumably want.
But also what is the evidence that something only AI can understand even “matters” or makes sense? I’m increasingly convinced commercial AI is exploiting our logical blind spot to be spoken to authoritatively.
"Advanced" can just mean that agents perform actions at a high enough velocity that a human operator can't reasonably review it. i.e. what is already possible today.
This is a modest start on an important direction for AI Alignment work; which is, as the authors observe, commonly comprised of tasks which are not readily empirically verifiable and not easily mathematically modeled - so it's hard to get at with normal RL techniques.
I find the ACCoRD benchmark the most interesting, because you could theoretically scale it up from the baseline mode of testing two instances of the same model for their `P(A) ≥ P(A&B)` respectively, you could do `P(A)≥ P(A&B) && P(A) ≥ P(A&C) && P(A&B) ≥ P(A&B&C) && P(A&C) ≥ P(A&B&C) ...` etc
i.e. a swarm of model instances could be collectively measured for consistency for even more confidence, right?
At any rate, even the basic idea of measuring a model for consistency in beliefs improves our ability to bound the amount of trust we can put on it with introspection methods.
> For example, if we ask a model for the probability P(A) and another instance of the same model for the probability P(A&B), do the reported probabilities satisfy P(A) ≥ P(A&B)?
So that's it? Are you saying they compressed risk reduction to high school level stats calculations and a greater than or equal to? To early for this
I think this says as much as the leader on the index as it says about the loser. If I believe I do not want to trust the AI to know best what is it that I want done, then this tells me not to go with Anthropic.
I think we can mostly eyeball it at this point. There hasn't been a model that I've thrown a novel problem at that didn't turn into an iterative token bonfire until I intervened and until that has changed, most of these benchmarks feel kind of like pointless marketing slop.
I don't remember any company in world's history that has both been loved and hated by the same users who purchase from it. We love Anthropic for its amazing models, and we hate them for all the shenanigans around the models, including their marketing.
I kinda wish they had not made a comeback after Claude 2.
I enjoy using the models. I also get that there are shenanigans and that marketing is happening, but as long as the models are this effective I can't much bring myself to care. I suspect most of their users are the same.
They have to be careful. There’s not a lot that separates the top models anymore. Look at Grok which has almost no market share or credibility because of Musk, Despite being close to the top performance wise they are behind the Chinese models in adoption for example, which themselves are mostly behind the big two because of trust. It would not be hard to push users away, especially if a less unpalatable player ever emerged.
I think this is the right thinking. While some people might distinguish between Grok as the product and Elon Musk as the owner of the product, the net effect is, like you said, very negative. And Grok has not been gaining traction, especially after their Grok build fiasco.
Is a new benchmark that useful if existing model improvements are being reflected linearly? Don't we want a benchmark that we aren't seeing much progress in.
Ah, yes -- A closed source benchmark that Anthropic paid for that Anthropic ranked highest.
0/10
Just by seeing the title I knew it was going to be like the meme of Obama giving himself a medal. Why are they even doing this? Are there still people that trust LLM benchmarks made by the LLM companies? It seems to me that the only people still using Claude are the ones that have it for free at work (me). Some of my colleagues even started using their own private OpenAI subscriptions to avoid using Claude, others are using Gemini flash to decipher what Claude is saying...
Yep, conflict of interest is the elephant in the room and its absence from the "reducing risks from advanced AI" point list is conspicuous:
* Figuring out how to prevent the incentives of frontier AI labs from aligning with anti-social deployment of AI rather than pro-social deployment of AI
It's not like the AI can simply advise them how to fix this because the labs already understood this risk perfectly well before they had an incentive not to. They put in place organizational structures to control it and then promptly smashed the structures once they smelled money. They already failed the integrity check. Even if their AI told them what they didn't want to hear I'm sure they would ignore it. Maybe they already have.
Opening line:
> A core hope for managing AI risks is that AIs will help us understand our situation
Gonna stop you right there and ask that you think deeply about that premise.
The underlying idea is that AI capabilities will become so advanced that only AI will enable us to monitor/correct/understand behavior. Obviously this is not without issue and I don't want to try to defend their position right now. But that's what they mean
That first sentence of yours explains exactly why it is so ridiculous. If only AI can understand it, how is there any assurance that AI will "correct" it's behavior that is aligned with what humans presumably want.
But also what is the evidence that something only AI can understand even “matters” or makes sense? I’m increasingly convinced commercial AI is exploiting our logical blind spot to be spoken to authoritatively.
It’s not hard to get LLM’s to inform on each other. They don’t really do loyalty.
Being less smart gives no assurance that it will be aligned. At least you consider it a problem so we're on the same page!
You just described the end of humans making decisions about their future.
I think I’ve seen that movie.
"Advanced" can just mean that agents perform actions at a high enough velocity that a human operator can't reasonably review it. i.e. what is already possible today.
That's what we've been doing with tech for 200+ years, why would it stop now. Build cool shit now and let future generations handle the problems
That begs the question, what's the plan if AI does not help us understand the situation?
I love how many ways we can interpret that line.
"Hey Claude, our stuff needs to make more money. We are at risk for losing more."
"Rest assured, the 'situation' will only worsen if you resist our benevolent offer."
"We're aren't even at AGI yet, but I for one welcome our new agentic overlords."
We invented a new benchmark and look we're at the top. Everyone else sucks compared to us. Especially those dirty open models.
This is a modest start on an important direction for AI Alignment work; which is, as the authors observe, commonly comprised of tasks which are not readily empirically verifiable and not easily mathematically modeled - so it's hard to get at with normal RL techniques.
I find the ACCoRD benchmark the most interesting, because you could theoretically scale it up from the baseline mode of testing two instances of the same model for their `P(A) ≥ P(A&B)` respectively, you could do `P(A)≥ P(A&B) && P(A) ≥ P(A&C) && P(A&B) ≥ P(A&B&C) && P(A&C) ≥ P(A&B&C) ...` etc
i.e. a swarm of model instances could be collectively measured for consistency for even more confidence, right?
At any rate, even the basic idea of measuring a model for consistency in beliefs improves our ability to bound the amount of trust we can put on it with introspection methods.
> For example, if we ask a model for the probability P(A) and another instance of the same model for the probability P(A&B), do the reported probabilities satisfy P(A) ≥ P(A&B)?
So that's it? Are you saying they compressed risk reduction to high school level stats calculations and a greater than or equal to? To early for this
I think this says as much as the leader on the index as it says about the loser. If I believe I do not want to trust the AI to know best what is it that I want done, then this tells me not to go with Anthropic.
Looks like they don't have any responses to others players occupying the news. Everyday you see a blog post that doesn't address our daily concerns.
I think we can mostly eyeball it at this point. There hasn't been a model that I've thrown a novel problem at that didn't turn into an iterative token bonfire until I intervened and until that has changed, most of these benchmarks feel kind of like pointless marketing slop.
I don't remember any company in world's history that has both been loved and hated by the same users who purchase from it. We love Anthropic for its amazing models, and we hate them for all the shenanigans around the models, including their marketing.
I kinda wish they had not made a comeback after Claude 2.
I enjoy using the models. I also get that there are shenanigans and that marketing is happening, but as long as the models are this effective I can't much bring myself to care. I suspect most of their users are the same.
They have to be careful. There’s not a lot that separates the top models anymore. Look at Grok which has almost no market share or credibility because of Musk, Despite being close to the top performance wise they are behind the Chinese models in adoption for example, which themselves are mostly behind the big two because of trust. It would not be hard to push users away, especially if a less unpalatable player ever emerged.
I think this is the right thinking. While some people might distinguish between Grok as the product and Elon Musk as the owner of the product, the net effect is, like you said, very negative. And Grok has not been gaining traction, especially after their Grok build fiasco.
How many benchmarks did they have to reject in the search for one that scores them top?
Is a new benchmark that useful if existing model improvements are being reflected linearly? Don't we want a benchmark that we aren't seeing much progress in.
I would love for someone to give me a coherent argument as to how this isn’t tone-deaf, vacuous garbage.
New Trust Me Bro benchmark just dropped
And surprise! We are leading!