Goto

Collaborating Authors

 Personal


TOD-ProcBench: Benchmarking Complex Instruction-Following in Task-Oriented Dialogues

arXiv.org Artificial Intelligence

In real-world task-oriented dialogue (TOD) settings, agents are required to strictly adhere to complex instructions while conducting multi-turn conversations with customers. These instructions are typically presented in natural language format and include general guidelines and step-by-step procedures with complex constraints. Existing TOD benchmarks often oversimplify the complex nature of these instructions by reducing them to simple schemas composed of intents, slots, and API call configurations. To address this gap and systematically benchmark LLMs' instruction-following capabilities, we propose TOD-ProcBench, a challenging benchmark featuring complex process instructions with intricate, fine-grained constraints that evaluates various LLMs' abilities to understand and follow instructions in multi-turn TODs. Our benchmark dataset comprises instruction documents derived from the high-quality ABCD dataset with corresponding conversations under human quality control. We formulate fine-grained constraints and action procedures as multi-level condition-action instruction statements. We design three tasks to comprehensively benchmark LLMs' complex instruction-following capabilities in multi-turn TODs. Task 1 evaluates how LLMs retrieve the most relevant statement from a complex instruction and predict the corresponding next action. In Task 2, we synthesize instruction-violating responses by injecting inconsistencies and manipulating the original instructions, and then we analyze how effectively LLMs can identify instruction-violating responses. Task 3 investigates LLMs' abilities in conditional generation of instruction-following responses based on the original complex instructions. Additionally, we conduct studies on the impact of multilingual settings and different instruction text formats on compliance performance. We release our benchmark under the Llama 3.3 Community License Agreement.


Can Artificial Intelligence Accelerate Technological Progress? Researchers' Perspectives on AI in Manufacturing and Materials Science

arXiv.org Artificial Intelligence

Applications of artificial intelligence or machine learning in research Modes of use Surrogate modeling for physics - based models Modeling of poorly understood phenomena Data preprocessing Large language model use Applications AI/ML as research tool Production process design, monitoring, & output prediction Part design & properties prediction Materials design & properties prediction AI/ML as research product Generative AI design tool for consumers Generic research tasks Large language models for coding Large language models for literature review Benefits of artificial intelligence or machine learning in research Reduction in accuracy/cost/speed trade - off in research, especially computer modeling Reduced computation time Replacing experimentation Reducing need for computationally intensive, physics - based models Saving research labor Exploring larger design spaces Address of previously unsolvable problems Model poorly understood relationships between variables Identify human - unidentifiable patterns or phenomena Downsides of artificial intelligence or machine learning in research Accuracy weaknesses Predict poorly outside regions of dense, high - quality training data Interpretability weaknesses Bounds of accuracy can be unclear Accuracy assessment can be difficult Long - run scientific progress concerns AI/ML cannot develop novel scientific theory AI/ML may bypass opportunities to identify empirical or theoretical novelties Resource issues Data acquisition and cleaning is time - intensive AI/ML models are computation - and energy - intensive to develop Inappropriate use issues Easy to over - trust May be inappropriately used to address problems soluble with simpler methods 8 Second, AI/ML models can be trained on input and output data for phenomena (e.g., complex production processes) which lack robust theoretical models, developing novel predictive capabilities in the absence of explicit, human - designed theory. This is somet imes referred to as "phenomenological modeling," as it attempts to model phenomena in the absence of mechanistic, explanatory understanding: [T]he first reason we choose to use AI is because we don't have a good model of what our system is. . . I get a bunch of data coming in and I have a bunch of sensor readings, you know. . . And I use the AI to map the bunch of sensor readings to the process health or process status or machine status that I have.


RescueLens: LLM-Powered Triage and Action on Volunteer Feedback for Food Rescue

arXiv.org Artificial Intelligence

Food rescue organizations simultaneously tackle food insecurity and waste by working with volunteers to redistribute food from donors who have excess to recipients who need it. Volunteer feedback allows food rescue organizations to identify issues early and ensure volunteer satisfaction. However, food rescue organizations monitor feedback manually, which can be cumbersome and labor-intensive, making it difficult to prioritize which issues are most important. In this work, we investigate how large language models (LLMs) assist food rescue organizers in understanding and taking action based on volunteer experiences. We work with 412 Food Rescue, a large food rescue organization based in Pittsburgh, Pennsylvania, to design RescueLens, an LLM-powered tool that automatically categorizes volunteer feedback, suggests donors and recipients to follow up with, and updates volunteer directions based on feedback. We evaluate the performance of RescueLens on an annotated dataset, and show that it can recover 96% of volunteer issues at 71% precision. Moreover, by ranking donors and recipients according to their rates of volunteer issues, RescueLens allows organizers to focus on 0.5% of donors responsible for more than 30% of volunteer issues. RescueLens is now deployed at 412 Food Rescue and through semi-structured interviews with organizers, we find that RescueLens streamlines the feedback process so organizers better allocate their time.


How Should the Law Treat Future AI Systems? Fictional Legal Personhood versus Legal Identity

arXiv.org Artificial Intelligence

The law draws a sharp distinction between objects and persons, and between two kinds of persons, the ''fictional'' kind (i.e. corporations), and the ''non-fictional'' kind (individual or ''natural'' persons). This paper will assess whether we maximize overall long-term legal coherence by (A) maintaining an object classification for all future AI systems, (B) creating fictional legal persons associated with suitably advanced, individuated AI systems (giving these fictional legal persons derogable rights and duties associated with certified groups of existing persons, potentially including free speech, contract rights, and standing to sue ''on behalf of'' the AI system), or (C) recognizing non-fictional legal personhood through legal identity for suitably advanced, individuated AI systems (recognizing them as entities meriting legal standing with non-derogable rights which for the human case include life, due process, habeas corpus, freedom from slavery, and freedom of conscience). We will clarify the meaning and implications of each option along the way, considering liability, copyright, family law, fundamental rights, civil rights, citizenship, and AI safety regulation. We will tentatively find that the non-fictional personhood approach may be best from a coherence perspective, for at least some advanced AI systems. An object approach may prove untenable for sufficiently humanoid advanced systems, though we suggest that it is adequate for currently existing systems as of 2025. While fictional personhood would resolve some coherence issues for future systems, it would create others and provide solutions that are neither durable nor fit for purpose. Finally, our review will suggest that ''hybrid'' approaches are likely to fail and lead to further incoherence: the choice between object, fictional person and non-fictional person is unavoidable.


Long-form factuality in large language models Jerry Wei 1 Chengrun Y ang 1 Xinying Song 1 Yifeng Lu

Neural Information Processing Systems

To benchmark a model's long-form factuality in open domains, we first use GPT -4 to generate LongFact, a prompt set comprising thousands of questions spanning 38 topics. We then propose that LLM agents can be used as automated evaluators for long-form factuality through a method which we call Search-Augmented Factuality Evaluator (SAFE).


In Alex Karp's World, Palantir Is the Underdog

WIRED

My parents didn't go to college, but his father was a pediatrician, Jewish American. His mother was an artist that still is an artist, and she's African American. So he is Black and Jewish parentage. He is dyslexic, and that's a big part of his identity. And when we talked about going to Central High School, which is kind of a magnet school, it's all academic, and it draws from all over the city.



Why quasicrystals shouldn't exist but are turning up in strange places

New Scientist

Why quasicrystals shouldn't exist but are turning up in strange places Matter with "forbidden" symmetries was once thought to be confined to lab experiments, but is now being found in some of the world's most extreme environments In autumn 1945, Lincoln LaPaz crouched over a patch of scorched ground in the Jornada del Muerto desert of New Mexico. LaPaz, an astronomer, was out hunting for meteorites. He had spotted something in the dust: a strange, glittering crust of blood-red glass. This was no meteorite, but it was striking enough that he held onto it. It wasn't until decades later that anyone would realise quite how special LaPaz's chance find was.


Londoners are baffled as a huge AI-generated Christmas mural appears over Cรดte Brasserie in Kingston - so, can you see what's wrong with it?

Daily Mail - Science & tech

Elon Musk caught playing with fire as he appears to makes explosive remark at Trump's Saudi banquet Clinton's private chat with'Hollywood' Gavin sets tongues wagging... and suddenly 2028 looks very different Wall Street hits'extreme fear' and stocks plunge. So we spoke to dozens of investment experts... and they all said exactly the same thing about your 401k: Read their urgent advice now Twist in cheerleader's mystery cruise ship death as FBI eyes shock suspect in criminal investigation Melania's subtle gesture to Saudi prince as she stuns in strapless green gown after Trump's extraordinary Oval Office defense sparked outrage NASA scientists are baffled to discover a rock on Mars that'doesn't belong there' This little-known skin condition ruined my life. It's not acne, eczema or even rosacea - but a combination of all three that appears out of nowhere and affects thousands. What really happened to Tati Westbrook: Her YouTube spat with James Charles backfired... then things took an even uglier turn. Gustav Klimt painting sells for $236.4 million as the most expensive piece of modern art ever sold at auction Carnage on America's roads as new deadly threat sparks widespread alarm: Read our full investigation Hakeem Jeffries becomes latest Democrat stung by Epstein files as he insists he'never met' billionaire Cristiano Ronaldo's touching moment with Barron Trump revealed as soccer star attends glitzy White House dinner'I can't listen to music any more.


Vaping Is 'Everywhere' in Schools--Sparking a Bathroom Surveillance Boom

WIRED

Schools in the US are installing vape-detection tech in bathrooms to thwart student nicotine and cannabis use. A new investigation reveals the impact of using spying to solve a problem. It was in physical education class when Laila Gutierrez swapped out self-harm for a new vice. The freshman from Phoenix had long struggled with depression and would cut her arms to feel something. The first drag from a friend's vape several years ago offered the shy teenager a new way to escape. She quit cutting but got hooked on nicotine. Her sadness got harder to carry after her uncle died, and she felt she couldn't turn to her grieving parents for comfort. Bumming fruity vapes at school became part of her routine. "I would ask my friends who had them, 'I'm going through a lot, can I use it?'" Gutierrez, now 18, told The 74. "Or'I failed my test and I feel like smoking would be better than cutting my wrists.'"