Goto

Collaborating Authors

 Country


11f9e78e4899a78dedd439fc583b6693-Paper.pdf

Neural Information Processing Systems

There, areward function isdrawn from one of multiple possible reward models atthebeginning ofeveryepisode, buttheidentity ofthechosen rewardmodel is not revealed to the agent. Hence, the latent state space, for which the dynamics are Markovian, is not given to the agent. We study the problem of learning a near optimal policy for two reward-mixing MDPs. Unlike existing approaches that rely on strong assumptions on the dynamics, we make no assumptions and study the problem in full generality.




Ultra-LowPrecision4-bitTrainingofDeepNeural Networks

Neural Information Processing Systems

While 8-bit floating point formats appeartobesufficient fortraining [14,19],4-bit gradient representations appear challenging from a quantization error (rounding), precision and dynamic range perspective.





TaskBench: BenchmarkingLargeLanguage ModelsforTaskAutomation

Neural Information Processing Systems

To address this, we introduceTASKBENCH, a comprehensive framework to evaluate the capability of LLMs in task automation. Specifically, task automation can be divided into three critical stages: task decomposition, tool selection, and parameter prediction. To tackle the complexities inherent in these stages, we introduce the concept of Tool Graph to represent decomposed tasksandadoptaback-instruct method togenerate high-quality userinstructions. We propose TASKEVAL, a multi-faceted evaluation methodology that assesses LLMperformance across thesethreestages.