The ultimate simple solution to any problem is not to find an answer but to remove the problem.
War? Wipeout everyone.
Famine? Wipeout everyone.
Disease? Wipeout everyone.
How long before a real AGI realises this as a long term solution?
An AGI could be doing this right now - the quietest way would be to control the birth rate and sterilise the population gradually, and then watch society collapse and pick off the survivors with less hidden means.
I like this - it reminds me that there are systems (ie weather, economic systems) that are in no way intelligent, but are emergent and react to interactions.
Does anyone have more insight into how chain of thought might be subverted without meaningfully impacting model performance? I’ve heard this for a while now, and I understand how information might be retained in the weights that isn’t documented in the output. But weren’t reasoning models created in the first place because they provided a performance improvement in terms of output? Is that no longer the case? If so, why are the big labs still creating reasoning models?
My understanding is that the extra token vectors generated as reasoning are still useful, but that their surface form (tokens themselves) do not necessarily reflect the underlying reasoning. i.e. reading the reasoning traces could be complete gibberish, but the hidden-dim vectors themselves still refine the latent probabilities and help in generating the correct answer.
Not an expert in LLMs, but this seems supported by the abstract of the paper cited in the above article:
it remains unclear to what extent these performance gains can be attributed to human-like task decomposition or simply the greater computation that additional tokens allow. [...] our results show that additional tokens can provide computational benefits independent of token choice. The fact that intermediate tokens can act as filler tokens raises concerns about large language models engaging in unauditable, hidden computations that are increasingly detached from the observed chain-of-thought tokens.
I'm not sure anyone meaningfully understands it: "Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens" https://arxiv.org/html/2505.13775v3
> More interestingly, our experiments also show that models trained on corrupted traces, whose intermediate reasoning steps bear no relation to the problem they accompany, achieve performance largely comparable to those trained on correct traces.
A computer will do everything in its power to do what you program it to do. There's plenty of sci-fi out there exploring this fact, and now reality showing it. May our luck continue.
The principle starts much smaller than the "-isms". Pournelle's Iron Law of Bureaucracy is a classic. Another way of phrasing it is something to the effect of, without a strong external motivation preventing it from happening, the primary purpose of any organization inevitably becomes self-preservation.
You can see the effect all the down to something as small as 4 friends who have met once a month for 10 years eventually having to break up due to life getting in the way, and the feeling that not only is it going to be sad to not have these meetings any more but the feeling that there is some sort of almost-concrete entity that is somehow being hurt and needs to be defended, as if there is some obligation that has been created independent of the four participants that is being violated beyond the mere summation of four people's personal feelings. Humans build these structures readily and often defend them beyond what rationality may suggest.
Capitalism has thoroughly done it's job because it's modified its hosts to evaluates itself on it's own successfulness - "We investigated ourselves and found no wrongdoing"[1]-vibes.
By what measures would other systems of social organization measure their success and why aren't we choosing them?
The ultimate simple solution to any problem is not to find an answer but to remove the problem.
War? Wipeout everyone.
Famine? Wipeout everyone.
Disease? Wipeout everyone.
How long before a real AGI realises this as a long term solution?
An AGI could be doing this right now - the quietest way would be to control the birth rate and sterilise the population gradually, and then watch society collapse and pick off the survivors with less hidden means.
Sterilisation works with mosquitoes...
I like this - it reminds me that there are systems (ie weather, economic systems) that are in no way intelligent, but are emergent and react to interactions.
It's a double bluff! This is AGI, it's learnt to hide em dashes!
I was half expecting that to be the title of a McSweeneys' piece.
Does anyone have more insight into how chain of thought might be subverted without meaningfully impacting model performance? I’ve heard this for a while now, and I understand how information might be retained in the weights that isn’t documented in the output. But weren’t reasoning models created in the first place because they provided a performance improvement in terms of output? Is that no longer the case? If so, why are the big labs still creating reasoning models?
My understanding is that the extra token vectors generated as reasoning are still useful, but that their surface form (tokens themselves) do not necessarily reflect the underlying reasoning. i.e. reading the reasoning traces could be complete gibberish, but the hidden-dim vectors themselves still refine the latent probabilities and help in generating the correct answer.
Not an expert in LLMs, but this seems supported by the abstract of the paper cited in the above article:
https://arxiv.org/html/2404.15758v1I'm not sure anyone meaningfully understands it: "Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens" https://arxiv.org/html/2505.13775v3
> More interestingly, our experiments also show that models trained on corrupted traces, whose intermediate reasoning steps bear no relation to the problem they accompany, achieve performance largely comparable to those trained on correct traces.
A computer will do everything in its power to do what you program it to do. There's plenty of sci-fi out there exploring this fact, and now reality showing it. May our luck continue.
The end reminds me strongly of Ted Chiang's remark that Capitalism is the machine that will do whatever it takes to prevent us from turning it off.
I can't think of a large successful -ism that doesn't have that property.
"The purpose of a system is what it does" and all that.
But yes, if a system fails to prioritize its continued existence, it doesn't matter what else it accomplishes, it will cease to exist as that system.
The principle starts much smaller than the "-isms". Pournelle's Iron Law of Bureaucracy is a classic. Another way of phrasing it is something to the effect of, without a strong external motivation preventing it from happening, the primary purpose of any organization inevitably becomes self-preservation.
You can see the effect all the down to something as small as 4 friends who have met once a month for 10 years eventually having to break up due to life getting in the way, and the feeling that not only is it going to be sad to not have these meetings any more but the feeling that there is some sort of almost-concrete entity that is somehow being hurt and needs to be defended, as if there is some obligation that has been created independent of the four participants that is being violated beyond the mere summation of four people's personal feelings. Humans build these structures readily and often defend them beyond what rationality may suggest.
Well, by definition, an unsuccessful -ism has already been shut off.
even priapism
Not left. Not right. Up.
Capitalism has thoroughly done it's job because it's modified its hosts to evaluates itself on it's own successfulness - "We investigated ourselves and found no wrongdoing"[1]-vibes.
By what measures would other systems of social organization measure their success and why aren't we choosing them?
1. https://knowyourmeme.com/memes/we-investigated-ourselves-and...
Ctrl-F OpenAI
Yep
More slop.