Goto

Collaborating Authors

 caveat


Reviews: Approximating Interactive Human Evaluation with Self-Play for Open-Domain Dialog Systems

Neural Information Processing Systems

This paper explores interesting directions, in particular 1) using interactive settings to evaluate a model rather than a single answer, and 2) combining different automated metrics in a weighted sums to approximate human evaluation (e.g., based on sentiment). Reviewers have raised crucial points, regarding gameability (so that using the metrics for training a model is tricky if not followed by a non-gameable evaluation), and lack of comparability between different self-play. It's indeed a much better evaluation setting if the system does not control both sides (e.g., models being matched to the same set of fixed models), so authors should definitely follow that direction. However, I expect this work would still be interesting to the dialog community: many of the diagnostic advantages of the model-talking-to-model setting remain, in practice, especially because the model is in fact not trained with the self-play objective, but that criterion is only used post hoc (so the system can't extensively exploit it during training). In practice, a lot of the problems of the generations of a given model already show up during self-play, and the reasonable worry raised by reviewers that the model could exploit the metric remains theoretical at the moment.


Beware of "Explanations" of AI

arXiv.org Artificial Intelligence

Understanding the decisions made and actions taken by increasingly complex AI system remains a key challenge. This has led to an expanding field of research in explainable artificial intelligence (XAI), highlighting the potential of explanations to enhance trust, support adoption, and meet regulatory standards. However, the question of what constitutes a "good" explanation is dependent on the goals, stakeholders, and context. At a high level, psychological insights such as the concept of mental model alignment can offer guidance, but success in practice is challenging due to social and technical factors. As a result of this ill-defined nature of the problem, explanations can be of poor quality (e.g. unfaithful, irrelevant, or incoherent), potentially leading to substantial risks. Instead of fostering trust and safety, poorly designed explanations can actually cause harm, including wrong decisions, privacy violations, manipulation, and even reduced AI adoption. Therefore, we caution stakeholders to beware of explanations of AI: while they can be vital, they are not automatically a remedy for transparency or responsible AI adoption, and their misuse or limitations can exacerbate harm. Attention to these caveats can help guide future research to improve the quality and impact of AI explanations.


Supercharging academic writing with generative AI: framework, techniques, and caveats

arXiv.org Artificial Intelligence

Academic writing is an indispensable yet laborious part of the research enterprise. This Perspective maps out principles and methods for using generative artificial intelligence (AI), specifically large language models (LLMs), to elevate the quality and efficiency of academic writing. We introduce a human-AI collaborative framework that delineates the rationale (why), process (how), and nature (what) of AI engagement in writing. The framework pinpoints both short-term and long-term reasons for engagement and their underlying mechanisms (e.g., cognitive offloading and imaginative stimulation). It reveals the role of AI throughout the writing process, conceptualized through a two-stage model for human-AI collaborative writing, and the nature of AI assistance in writing, represented through a model of writing-assistance types and levels. Building on this framework, we describe effective prompting techniques for incorporating AI into the writing routine (outlining, drafting, and editing) as well as strategies for maintaining rigorous scholarship, adhering to varied journal policies, and avoiding overreliance on AI. Ultimately, the prudent integration of AI into academic writing can ease the communication burden, empower authors, accelerate discovery, and promote diversity in science.


Caveats on the first-generation da Vinci Research Kit: latent technical constraints and essential calibrations

arXiv.org Artificial Intelligence

Telesurgical robotic systems provide a well established form of assistance in the operating theater, with evidence of growing uptake in recent years. Until now, the da Vinci surgical system (Intuitive Surgical Inc, Sunnyvale, California) has been the most widely adopted robot of this kind, with more than 6,700 systems in current clinical use worldwide [1]. To accelerate research on robotic-assisted surgery, the retired first-generation da Vinci robots have been redeployed for research use as "da Vinci Research Kits" (dVRKs), which have been distributed to research institutions around the world to support both training and research in the sector. In the past ten years, a great amount of research on the dVRK has been carried out across a vast range of research topics. During this extensive and distributed process, common technical issues have been identified that are buried deep within the dVRK research and development architecture, and were found to be common among dVRK user feedback, regardless of the breadth and disparity of research directions identified. This paper gathers and analyzes the most significant of these, with a focus on the technical constraints of the first-generation dVRK, which both existing and prospective users should be aware of before embarking onto dVRK-related research. The hope is that this review will aid users in identifying and addressing common limitations of the systems promptly, thus helping to accelerate progress in the field.


Google Bard AI hands-on: A work in progress with plenty of caveats

Engadget

Google has made Bard more widely available to users in the US and the UK today, and I have been spending some time with the company's chatbot to see how its generative AI compares to ChatGPT and Bing AI. Like we saw in the screenshots Google provided with today's announcement, the interface here is very similar to Bing AI in that there is a wide text input at the bottom of the screen and a dialogue-based layout. But there are a few key differences between Google's and Microsoft's offerings. With Bing AI, you'll have to either hit Chat or scroll up from search results to get to the conversation page, whereas you don't have to do that for the Bard website. Microsoft has a broom icon to the left of the input bar to clear the slate and start a new topic, while Google has a column on the left with options for "Reset chat," "Bard Activity," "FAQ and "Help & Support." It's also worth noting the language Google painstakingly uses here. Once I navigated to the website, I was greeted with an ...


Sony PS5 update enables external storage for next-gen games, with a caveat

Washington Post - Technology News

Sony has stated previously that the console will support additional storage via internal M.2 drives in the future. As such, this update is something of a stopgap that lessens, but doesn't eliminate, the inconvenience of the PS5′s limited storage space. Some of the console's early titles have already run into storage issues, with "Call of Duty: Black Ops Cold War" requiring over 200 GBs of space when players also include the "Warzone" battle royale mode. Other titles, like "Hitman 3" and "Destiny 2," eat up over 100 GBs, according to a March article from Game Rant.


AI talent appears open to working on defense -- with caveats

#artificialintelligence

A new survey offers some evidence that most artificial intelligence experts are positive or neutral when it comes to working with the Pentagon on AI-enabled projects. Why it matters: Employee concerns have led some tech companies to pull back from working on defense-related projects in the past, but for many in the AI world, the chance to work on intellectually challenging projects -- and the Pentagon's not insignificant budget -- seems too good to pass up. What's happening: In a report released earlier this week -- with the memorable title "'Cool Projects' or'Expanding the Efficiency of the Murderous American War Machine?'" -- researchers at CSET surveyed 160 AI professionals about their attitudes toward working on DoD projects. Flashback: In 2018, after thousands of employees signed a protest letter, Google CEO Sundar Pichai pulled out of a contract to work on Project Maven, a DoD-funded pilot AI program to develop computer vision algorithms. Yes, but: That narrative was "overblown," argues New America's Singer.


Those Three Clever Dogs Trained To Drive A Car Provide Valuable Lessons For AI Self-Driving Cars

#artificialintelligence

Perhaps this dog would prefer driving the car, just like three dogs that were trained to do so. We've all seen dogs traveling in cars, including how they like to peek out an open window and enjoy the fur-fluffing breeze and dwell in the cacophony of scents that blow along in the flavorful wind. Dogs have also frequently been used as living props in commercials for cars, pretending in some cases to drive a car, such as the Subaru "Barkleys" advertising campaign that initially launched on TV in 2018 and continued in 2019, proclaiming that Subaru cars were "officially" dog tested and dog approved. What you might not know or might not remember is that there were three dogs that were trained on driving a car and had their moment of unveiling in December of 2012 when they were showcased by driving a car on an outdoor track (the YouTube posted video has amassed millions of views). Yes, three dogs named Monty, Ginny, and Porter were destined to become the first true car drivers on behalf of the entire canine family. Monty at the time was an 18-month-old giant schnauzer cross, while the slightly younger Ginny at one year of age was a beardie whippet cross, and Porter was a youthful 10-month-old beardie.


Ensemble Methods for Decision Trees

#artificialintelligence

Decision Trees are popular Machine Learning algorithms used for both regression and classification tasks. Their popularity mainly arises from their interpretability and representability, as they mimic the way the human brain takes decisions. However, to be interpretable, they pay a price in terms of prediction accuracy. To overcome this caveat, some techniques have been developed, with the goal of creating strong and robust models starting from'poor' models. Those techniques are known as'ensemble' methods and, in this article, I'm going to talk about three of them: Bagging, Random Forest and Boosting.


Voices in AI – Bonus: A Conversation with Hilary Mason

#artificialintelligence

Today's leading minds talk AI with host Byron Reese Listen to this episode or read the full transcript at www.VoicesinAI.com Byron Reese: This is Voices in AI, brought to you by Gigaom and I am Byron Reese. Today, our guest is Hilary Mason. She is the GM of Machine Learning at Cloudera, and the founder and CEO of Fast Forward Labs, and the Data Scientist in residence at Accel Partners, and a member of the Board of Directors at the Anita Borg Institute for Women in Technology, and the co-founder of hackNY.org. That's as far down as it would let me read in her LinkedIn profile, but I've a feeling if I'd clicked that'More' button, there would be a lot more.