Goto

Collaborating Authors

 Technology





. Figure 1 m n 100 1000 10 29 4 s 33 6 s 50 8 1 min 9 1 min 100 15 1 min 24 2 min Table 2: Time to reach relative improvement 10

Neural Information Processing Systems

We thank the reviewers for their comments. We then address reviewer's comments individually (due to space limits please zoom in the tiny figures). For [18] we used Alg. 2 We thank the reviewer for the additional reference, which we will add to the paper. Gradient Descent) applied in parallel to multiple starting points. We thank R2 for the reference "Entropic regularization of continuous optimal transport problems".



Text-AwareDiffusionforPolicyLearning

Neural Information Processing Systems

Training an agent to achieve particular goals or perform desired behaviors is often accomplished through reinforcement learning, especially in the absence of expert demonstrations. However, supporting novel goals or behaviors through reinforcement learning requires the ad-hoc design of appropriate reward functions, which quickly becomes intractable. Toaddress thischallenge, wepropose Text-AwareDiffusion forPolicyLearning (TADPoLe), which uses apretrained, frozen text-conditioned diffusion model to compute dense zero-shot reward signals for text-aligned policy learning.





Beyond Online Balanced Descent: An Optimal Algorithm for Smoothed Online Optimization

Neural Information Processing Systems

Weproveanewlower bound onthe competitive ratio of any online algorithm in the setting where the costs aremstrongly convex and the movement costs are the squared`2 norm. This lower bound shows that no algorithm can achieveacompetitiveratio that iso(m 1/2) asmtendstozero.