It might be beneficial while not being optimal on its own.
The obvious example is if it has different behaviour around local minima, it could be an altenate pathway out.
I have often wondered if doing training with radically different aproaches for the first few iterarions would avoid any method specific artifacts before the weights had time to denoise.
New paper by Sakana.ai [1]
[1]: https://arxiv.org/abs/2605.31022
~90% accuracy on MNIST. Sigh.
How does it do on CIFAR-10, or even better, ImageNet?
It might be beneficial while not being optimal on its own.
The obvious example is if it has different behaviour around local minima, it could be an altenate pathway out.
I have often wondered if doing training with radically different aproaches for the first few iterarions would avoid any method specific artifacts before the weights had time to denoise.
Their image classification benchmarks include both: https://pub.sakana.ai/pc-alm/assets/figures/benchmark_accura...