Follow-up on yesterday's hot-take: there are a lot of noise, and you should be suspicious of any claims such as 4x, 8x, 10x speed-up over
@awnihannun's MLX or
@ggerganov's llama.cpp with "custom / model-specific inference engines" on Apple's hardware.