Everyone will quote 2nm. The number that matters for local AI is this one:
50% more memory bandwidth than A19 Pro.
Decode is bandwidth-bound. You don't get 50% more tokens/sec from a faster core.
You get it from moving weights faster. This is the biggest single-generation bandwidth jump I've seen in an iPhone.