A Conditional Randomization Test for Sparse Logistic Regression in High-Dimension

May-30-2025, 00:14:00 GMT–Neural Information Processing Systems

Identifying the relevant variables for a classification model with correct confidence levels is a central but difficult task in high-dimension. Despite the core role of sparse logistic regression in statistics and machine learning, it still lacks a good solution for accurate inference in the regime where the number of features p is as large as or larger than the number of samples n. Here we tackle this problem by improving the Conditional Randomization Test (CRT). The original CRT algorithm shows promise as a way to output p-values while making few assumptions on the distribution of the test statistics. As it comes with a prohibitive computational cost even in mildly high-dimensional problems, faster solutions based on distillation have been proposed. Yet, they rely on unrealistic hypotheses and result in low-power solutions.

artificial intelligence, crt-logit, machine learning, (15 more...)

Neural Information Processing Systems

May-30-2025, 00:14:00 GMT

Conferences PDF

Add feedback

Country:
- Europe > United Kingdom > Scotland (0.14)

Genre:
- Research Report > Experimental Study (1.00)

Industry:
- Health & Medicine
  - Health Care Technology (1.00)
  - Pharmaceuticals & Biotechnology (1.00)
  - Therapeutic Area
    - Neurology (1.00)
    - Oncology (1.00)

Technology:
- Information Technology > Artificial Intelligence > Machine Learning > Statistical Learning > Regression (0.72)