Beyond 2:4: exploring V:N:M sparsity for efficient transformer inference on GPUs

Open in new window