left-hand side
Time-adaptive functional Gaussian Process regression
Ruiz-Medina, MD, Madrid, AE, Torres-Signes, A, Angulo, JM
This paper proposes a new formulation of functional Gaussian Process regression in manifolds, based on an Empirical Bayes approach, in the spatiotemporal random field context. We apply the machinery of tight Gaussian measures in separable Hilbert spaces, exploiting the invariance property of covariance kernels under the group of isometries of the manifold. The identification of these measures with infinite-product Gaussian measures is then obtained via the eigenfunctions of the Laplace-Beltrami operator on the manifold. The involved time-varying angular spectra constitute the key tool for dimension reduction in the implementation of this regression approach, adopting a suitable truncation scheme depending on the functional sample size. The simulation study and synthetic data application undertaken illustrate the finite sample and asymptotic properties of the proposed functional regression predictor.
We thank all four reviewers for their thoughtful reviews, and are happy that they value the contribution of a new
's suggestion to draw the external nodes on the left-hand side differs from But this wouldn't be as flexible as we'd like; for example, we'd like to query a HMM for the But conjunction is an operation on FGGs, not factor graphs, so at the time of conjunction, no renaming has taken place. We agree that the notation should be improved and will think about how to do so. Lemma 15 does not change the generated graphs and cannot change their treewidth. We mean that an FGG can't generate the
A Experimental Details
We prove this by contradiction. The original problem in Eq. (2) is now equivalently reduced following problem because r Namely, it is the solution of a ridgeless linear regression problem. In the second case, one needs to discern the saddle points from the global minima. The objective (9) can be upper-bounded by Eq. (9) γ We see that the above inequality is equivalent to Eq. (9) γ B.5.1 Proposition 4 Proposition 4. Any global minimum of Eq. (9) is of the form U = b By the induction assumption, the global minimum of this problem takes the form of Eq. B.6 Lemma 3 Lemma 3. At any global minimum of Eq. (9), let b ( D 1) We first apply Lemma 3 to determine the condition for the nontrivial solution to exist.
LogicQA: Logical Anomaly Detection with Vision Language Model Generated Questions
Kwon, Yejin, Moon, Daeun, Oh, Youngje, Yoon, Hyunsoo
Anomaly Detection (AD) focuses on detecting samples that differ from the standard pattern, making it a vital tool in process control. Logical anomalies may appear visually normal yet violate predefined constraints on object presence, arrangement, or quantity, depending on reasoning and explainability. We introduce LogicQA, a framework that enhances AD by providing industrial operators with explanations for logical anomalies. LogicQA compiles automatically generated questions into a checklist and collects responses to identify violations of logical constraints. LogicQA is training-free, annotation-free, and operates in a few-shot setting. We achieve state-of-the-art (SOTA) Logical AD performance on public benchmarks, MVTec LOCO AD, with an AUROC of 87.6 percent and an F1-max of 87.0 percent along with the explanations of anomalies. Also, our approach has shown outstanding performance on semiconductor SEM corporate data, further validating its effectiveness in industrial applications.
Platforms for Efficient and Incentive-Aware Collaboration
Haghtalab, Nika, Qiao, Mingda, Yang, Kunhe
Collaboration is crucial for reaching collective goals. However, its effectiveness is often undermined by the strategic behavior of individual agents -- a fact that is captured by a high Price of Stability (PoS) in recent literature [Blum et al., 2021]. Implicit in the traditional PoS analysis is the assumption that agents have full knowledge of how their tasks relate to one another. We offer a new perspective on bringing about efficient collaboration among strategic agents using information design. Inspired by the growing importance of collaboration in machine learning (such as platforms for collaborative federated learning and data cooperatives), we propose a framework where the platform has more information about how the agents' tasks relate to each other than the agents themselves. We characterize how and to what degree such platforms can leverage their information advantage to steer strategic agents toward efficient collaboration. Concretely, we consider collaboration networks where each node is a task type held by one agent, and each task benefits from contributions made in their inclusive neighborhood of tasks. This network structure is known to the agents and the platform, but only the platform knows each agent's real location -- from the agents' perspective, their location is determined by a random permutation. We employ private Bayesian persuasion and design two families of persuasive signaling schemes that the platform can use to ensure a small total workload when agents follow the signal. The first family aims to achieve the minmax optimal approximation ratio compared to the optimal collaboration, which is shown to be $\Theta(\sqrt{n})$ for unit-weight graphs, $\Theta(n^{2/3})$ for graphs with constant minimum edge weights, and $O(n^{3/4})$ for general weighted graphs. The second family ensures per-instance strict improvement compared to full information disclosure.
Greedy Algorithm for Inference of Decision Trees from Decision Rule Systems
Durdymyradov, Kerven, Moshkov, Mikhail
Decision trees [3, 4, 8, 31, 34, 40] and systems of decision rules [6, 7, 11, 14, 33, 34, 35, 36] are widely used as classifiers, knowledge representation tools, and algorithms. They are known for their interpretability in data analysis [10, 15, 23, 41]. Investigating the relationship between these two models is an important task in computer science. Converting decision trees into decision rule systems is a well-known and simple process [37, 38, 39]. This paper focuses on the inverse transformation problem, which is not trivial. The research related to this problem encompasses several directions: Two-stage construction of decision trees. This approach involves building decision rules based on input data, followed by the construction of decision trees or decision structures (which are generalizations of decision trees) using the generated rules. The benefits of this two-stage construction method are explained in [1, 2, 17, 18, 19, 20, 21, 22, 42].
Compression, Generalization and Learning
Campi, Marco C., Garatti, Simone
A compression function is a map that slims down an observational set into a subset of reduced size, while preserving its informational content. In multiple applications, the condition that one new observation makes the compressed set change is interpreted that this observation brings in extra information and, in learning theory, this corresponds to misclassification, or misprediction. In this paper, we lay the foundations of a new theory that allows one to keep control on the probability of change of compression (which maps into the statistical "risk" in learning applications). Under suitable conditions, the cardinality of the compressed set is shown to be a consistent estimator of the probability of change of compression (without any upper limit on the size of the compressed set); moreover, unprecedentedly tight finite-sample bounds to evaluate the probability of change of compression are obtained under a generally applicable condition of preference. All results are usable in a fully agnostic setup, i.e., without requiring any a priori knowledge on the probability distribution of the observations. Not only these results offer a valid support to develop trust in observation-driven methodologies, they also play a fundamental role in learning techniques as a tool for hyper-parameter tuning.
Observable adjustments in single-index models for regularized M-estimators
We consider observations $(X,y)$ from single index models with unknown link function, Gaussian covariates and a regularized M-estimator $\hat\beta$ constructed from convex loss function and regularizer. In the regime where sample size $n$ and dimension $p$ are both increasing such that $p/n$ has a finite limit, the behavior of the empirical distribution of $\hat\beta$ and the predicted values $X\hat\beta$ has been previously characterized in a number of models: The empirical distributions are known to converge to proximal operators of the loss and penalty in a related Gaussian sequence model, which captures the interplay between ratio $p/n$, loss, regularization and the data generating process. This connection between$(\hat\beta,X\hat\beta)$ and the corresponding proximal operators require solving fixed-point equations that typically involve unobservable quantities such as the prior distribution on the index or the link function. This paper develops a different theory to describe the empirical distribution of $\hat\beta$ and $X\hat\beta$: Approximations of $(\hat\beta,X\hat\beta)$ in terms of proximal operators are provided that only involve observable adjustments. These proposed observable adjustments are data-driven, e.g., do not require prior knowledge of the index or the link function. These new adjustments yield confidence intervals for individual components of the index, as well as estimators of the correlation of $\hat\beta$ with the index. The interplay between loss, regularization and the model is thus captured in a data-driven manner, without solving the fixed-point equations studied in previous works. The results apply to both strongly convex regularizers and unregularized M-estimation. Simulations are provided for the square and logistic loss in single index models including logistic regression and 1-bit compressed sensing with 20\% corrupted bits.