New in
@berkeley_ai blog. Cao et al. create new kernel optimizer K-search and use it to auto adapt CUDA to MLX. Their attention translation measures 0.97x speed to Appleโs native implementation. Translate your code with their OSS tool. Available now.
๐ป๐