Register and share your invite link to earn from video plays and referrals.

Search results for mha 
mha  community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including mha 
#MHA# issues notification: #Haldia# port in #PurbaMedinipore# is now a designated #immigration# check post. Foreigners can now enter & exit from Haldia Seaport.
wabbit season!.... whatever happen to MHA and its super toxic fandom? 🤔🤔🤔
0
29
3.6K
213
Forward to community
Union Home Minister Amit Shah will inaugurate 3rd National Conference of Anti-Narcotics Task Force (ANTF) heads of states and Union Territories in New Delhi on September 22: MHA
New NanoGPT Speedrun WR at 76.3 (-2.9s) from @Lisennlp , with a radical boost to the final attention layer that drops step count by 7%, called lightweight Dynamically Composable MHA. A secondary local window (112 tokens) is recomputed on the same q/k. After softmax it mixes the per-head attn maps into one unified map, then applies the shared map to each head's values with a per-head scale. Every mix and scale operation is computed per-token from a MUDD MLP. All in a custom kernel. Takes inspiration from Dynamically Composable MHA
Show more
After reading up a bit on ML research post transformer era, I was upset that it seems to have converged on hyper-optimizing matmul-based algorithms: (MHA, MQA, MLA, SWA, DSA, GQA, SWA-GQA, ABCDA [only one of these is made up]). Surely, an algorithm that is not Attention based is sitting there waiting to be discovered. > the researchers are just being lazy but this is a stupid conclusion. How can you blame researchers, when the hardware they train on is optimized for matmuls (tensor cores / systolic arrays). Any algorithm not a matmul is literally bound to die, even if it's twice as good as attention. Add compute constraints, you have to be crazy to research any direction not attention based (basically @sarahookr 's hardware lottery essay) We talk about hardware-software co-design in inference, but it seems that, to get to the next leap in research, we'll need hardware-research co-design. At first, it seems this will never happen, given typical multi-year hardware tape-out constraints. But then you look at @OpenAI. 9 month tape-out. Better "training" and serving . Why fab your own chip if it's just going to be systolic-array based? Why not just buy Nvidia? > "But Nvidia GPUs are scarce" Then buy TPUs/AMD/Qualcom/Cerebras. Sure the software is not that good, but if you're OAi, you can hire an army of engineers to unlock the full capability. Either they moved away from attention and have a new algorithm they needed their own chips to train it on (unlikely given that a 9-month tape-out with a TPU vendor implies reusing IP)...or research is dead and we're never escaping attention / matmul based algo.
Show more