related work
A Further Related Work on Nonsmooth Nonconvex Optimization
To appreciate the difficulty and the broad scope of the research agenda in nonsmooth nonconvex optimization, we start by describing the existing relevant literature. First, the existing work is mostly devoted to establishing the asymptotic convergence properties of various optimization algorithms, including gradient sampling (GS) methods [16-18, 57, 19], bundle methods [56, 40] and subgradient methods [8, 65, 30, 28, 12]. More specifically, Burke et al. [16] provided a systematic investigation of approximating the Clarke subdifferential through random sampling and proposed a gradient bundle method [17]--the precursor of GS methods--for optimizing a nonconvex, nonsmooth and non-Lipschitz function. Later, Burke et al. [18] and Kiwiel [57] proposed the GS methods by incorporating key modifications into the algorithmic scheme in Burke et al. [17] and proved that every cluster point of the iterates generated by GS methods is a Clarke stationary point. For an overview of GS methods, we refer to Burke et al. [19].
A Related Work
For instance, one such notion is'unawareness', which necessitates Additionally, preference-based fairness argues that an algorithm's design should not be solely determined by its creators or regulators but should also incorporate the preferences of those directly A myriad of techniques exist to construct fair models using counterfactual inference. Theorem 2. Assume that R has been generated using Algorithm 2. We have, Pr(R We consider a causal graph shown in Figure 6. The counterfactual data ˇ X were computed by substituting A in the structural function with ˇ A . We implemented our method and the baseline methods as described in Section 5 (since there is no difference between observed data and factual data in this scenario, we have no ICA baseline here). For the CR method, we set the weight of the fairness regularization term as 0.05.
A Related Work .
Semantic IDs created using an auto-encoder (RQ-V AE [40, 21]) for retrieval models. We refer to V ector Quantization as the process of converting a high-dimensional vector into a low-dimensional tuple of codewords. We discuss this technique in more detail in Subsection 3.1. We use users' review history During training, we limit the number of items in a user's history to 20. The results for this dataset are reported in Table 7 as the row'P5'.
Supplementary Material for " Variational Policy Gradient Method for Reinforcement Learning with General Utilities " A Related Work
We provide a more extension discussion for the context of this work. Firstly, when closed-form expressions for the optimizer of a function are unavailable, solving optimization problems requires iterative schemes such as gradient ascent [31]. Their convergence to global extrema is predicated on concavity and the tractability of computing ascent directions. When the objective takes the form of an expected value of a function parameterized by a random variable, stochastic approximations are required [36, 24]. The PG Theorem mentioned above gives a specific form for obtaining ascent directions with respect to a parameterized family of stationary policies via trajectories in a Markov decision process, when the objective is the expected cumulative return [44], which gives rise to the REINFORCE algorithm.