When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity

Open in new window