Beyond 2:4: exploring V:N:M sparsity for efficient transformer inference on GPUs