How does this relate to the more recent work on the M4 ANE found at https://maderix.github.io/articles/ ? Does the M4 and later ANE expose any additional capabilities, or is it just a higher-performance iteration of the same thing?
As an aside, the introduction to this article seems to conflate the ANE with the Neural Accelerators (NAX) found in the M5+ (and A-series equivalents) GPUs. These are very different things, and Apple is still working on the ANE - the M6 and A20 will apparently feature doubled ANE blocks.
This isn't ai slop. It's fascinating and well written.
But I learned something really basic - i didn't know that the ANE (and the data pipeline around it) was designed for CNN rather than transformers. It's always been an open loop in my head, wondering why the ANE was less impactful than i understood it should be.
A lot of neural engine, particularly in the embedded domain (ARM/RISC MCUs) have the same problem. Designing other models means on top of this means a lot of profiling to get convolution blocks right to get good speedups. (We optimized this in the past e.g. using Neural Architecture Search on super networks)
Also, just imagine being the group at Apple responsible for designing this section of the chip, starting probably almost a decade back – under the constant uncertainty of not knowing what direction ML workloads would develop in…
How does this relate to the more recent work on the M4 ANE found at https://maderix.github.io/articles/ ? Does the M4 and later ANE expose any additional capabilities, or is it just a higher-performance iteration of the same thing?
As an aside, the introduction to this article seems to conflate the ANE with the Neural Accelerators (NAX) found in the M5+ (and A-series equivalents) GPUs. These are very different things, and Apple is still working on the ANE - the M6 and A20 will apparently feature doubled ANE blocks.
This isn't ai slop. It's fascinating and well written.
But I learned something really basic - i didn't know that the ANE (and the data pipeline around it) was designed for CNN rather than transformers. It's always been an open loop in my head, wondering why the ANE was less impactful than i understood it should be.
A lot of neural engine, particularly in the embedded domain (ARM/RISC MCUs) have the same problem. Designing other models means on top of this means a lot of profiling to get convolution blocks right to get good speedups. (We optimized this in the past e.g. using Neural Architecture Search on super networks)
> But I learned something really basic
Same for me!
Also, just imagine being the group at Apple responsible for designing this section of the chip, starting probably almost a decade back – under the constant uncertainty of not knowing what direction ML workloads would develop in…
> what workloads it was designed for and accels at.
excels!
Could've been a pun as a neural processor is an accelerator, so it 'accels' at machine learning tasks!
At least we know it wasn't written by a bot