there's a lot of juice to be squeezed from innovation and experimentation like this
stacked inference optimizations can deliver 4–6× speedups, or 2–4× with fixed hardware and GPU count.
@waterloo_intern explains -
think this is correct paper?
@the_joshua_hill