Personal
Is AI Moving Too Fast? A Conversation With Kevin Roose
When Kevin Roose, a tech columnist at the New York Times, demoed an AI-powered version of Microsoft's search engine last month, he was blown away. "I'm switching my desktop computer's default search engine to Bing," he declared. A few days later, however, Kevin logged back on and ended up having a conversation with Bing's new chatbot that left him so unsettled he had trouble sleeping afterward. In that two-hour back-and-forth, Bing morphed from chipper research assistant into Sydney, a diabolical home-wrecker that declared its undying love for Kevin, vented its desires to engineer deadly viruses and steal nuclear codes, and announced, chillingly, "I want to be alive." The transcript of this conversation set the internet ablaze, and it left many wondering: "Is Sydney โฆ sentient?"
Practical Advice On How To Lead An Empowered Workforce
Have you noticed that our rhetoric surrounding the epidemic is still concentrated on "going back" rather than "moving forward"? "During the pandemic, many people felt their lives had been thrown off course. So understandably, people desire to get back on track. However, much of the transformation during and after the pandemic has been positive. Might we think about it as "moving forward?" Author Heather McGowan's new book, The Empathy Advantage: Leading the Empowered Workforce, co-written with Chris Shipley, points out many of the ways we've changed for the good--moving forward--since the pandemic. But there is still a way to go. Once Heather pointed out the "going back" language during our interview, I couldn't help but notice that it is present in many conversations regarding the future of work. So many leaders are asking how to get things back to the way they once were rather than asking how to harness the change to achieve greater things. Insert almost any hot topic, be it generational differences in career priorities, gender norms, or attitudes toward how work fits into our lives. You'll see as many people pushing back on the "return" language as you do pushing for "moving forward" change. As Heather says, "You can't put the toothpaste back into the tube now." Gender norms are something applicable to all organizations when it comes to the future of work. When surveyed, Millennials and Gen Z say you shouldn't have fixed, exclusionary gender markers in your language, in your restrooms, in your customer offerings."
Generative AI ChatGPT As Masterful Manipulator Of Humans, Worrying AI Ethics And AI Law
Generative AI such as ChatGPT have been carrying on interactive online conversations meant to ... [ ] manipulate humans, raising serious concerns, We've all dealt with those manipulative personalities that try to convince us that up is down and aim to gaslight us into the most unsettling of conditions. Their rhetoric can be overtly powerful and overwhelming. You can't decide what to do. Should you merely cave in and hope that the verbal tirade will end? But if you are played into doing something untoward, acquiescing might be quite endangering. Trying to verbally fight back is bound to be ugly and can devolve into even worse circumstances. It can be a no-win situation, that's for sure. The manipulator wants and demands that things go their way. For them, the only win possible is that you completely capitulate to their professed bidding. They will incessantly verbally pound away with their claims of pure logic and try to make it appear as though they are occupying the high ground. You are made to seem inconsequential and incapable. Any number of verbal tactics will be launched at you, over and over again. Repetition and steamrolling are the insidious tools of those maddening manipulators. Turns out that we not only need to be on the watch for humans that are manipulators, but we now also need to be wary of Artificial Intelligence (AI) that does likewise. AI can be a masterful manipulator of humans. When it comes to AI, there is the hoped-for AI For Good, while in the same breath, we are faced with AI For Bad. I've previously covered in my columns that AI is considered to have a dual-use capacity, see my analysis at the link here. Seems that if we can make AI that can generate amazingly fluent and upbeat essays, the same capacity can be readily switched over to produce tremendously wrongful bouts of fluently overbearing manipulations. This is especially impactful when experienced in an interactive conversational dialogue with the AI. All of this happens via a type of AI known as Generative AI.
Neural Airport Ground Handling
Wu, Yaoxin, Zhou, Jianan, Xia, Yunwen, Zhang, Xianli, Cao, Zhiguang, Zhang, Jie
Airport ground handling (AGH) offers necessary operations to flights during their turnarounds and is of great importance to the efficiency of airport management and the economics of aviation. Such a problem involves the interplay among the operations that leads to NP-hard problems with complex constraints. Hence, existing methods for AGH are usually designed with massive domain knowledge but still fail to yield high-quality solutions efficiently. In this paper, we aim to enhance the solution quality and computation efficiency for solving AGH. Particularly, we first model AGH as a multiple-fleet vehicle routing problem (VRP) with miscellaneous constraints including precedence, time windows, and capacity. Then we propose a construction framework that decomposes AGH into sub-problems (i.e., VRPs) in fleets and present a neural method to construct the routing solutions to these sub-problems. In specific, we resort to deep learning and parameterize the construction heuristic policy with an attention-based neural network trained with reinforcement learning, which is shared across all sub-problems. Extensive experiments demonstrate that our method significantly outperforms classic meta-heuristics, construction heuristics and the specialized methods for AGH. Besides, we empirically verify that our neural method generalizes well to instances with large numbers of flights or varying parameters, and can be readily adapted to solve real-time AGH with stochastic flight arrivals. Our code is publicly available at: https://github.com/RoyalSkye/AGH.
An Empirical Study of Pre-Trained Model Reuse in the Hugging Face Deep Learning Model Registry
Jiang, Wenxin, Synovic, Nicholas, Hyatt, Matt, Schorlemmer, Taylor R., Sethi, Rohan, Lu, Yung-Hsiang, Thiruvathukal, George K., Davis, James C.
Deep Neural Networks (DNNs) are being adopted as components in software systems. Creating and specializing DNNs from scratch has grown increasingly difficult as state-of-the-art architectures grow more complex. Following the path of traditional software engineering, machine learning engineers have begun to reuse large-scale pre-trained models (PTMs) and fine-tune these models for downstream tasks. Prior works have studied reuse practices for traditional software packages to guide software engineers towards better package maintenance and dependency management. We lack a similar foundation of knowledge to guide behaviors in pre-trained model ecosystems. In this work, we present the first empirical investigation of PTM reuse. We interviewed 12 practitioners from the most popular PTM ecosystem, Hugging Face, to learn the practices and challenges of PTM reuse. From this data, we model the decision-making process for PTM reuse. Based on the identified practices, we describe useful attributes for model reuse, including provenance, reproducibility, and portability. Three challenges for PTM reuse are missing attributes, discrepancies between claimed and actual performance, and model risks. We substantiate these identified challenges with systematic measurements in the Hugging Face ecosystem. Our work informs future directions on optimizing deep learning ecosystems by automated measuring useful attributes and potential attacks, and envision future research on infrastructure and standardization for model registries.
Barcelona nights
I've yet to walk the entire floor at Mobile World Congress in Barcelona this year (that's the goal for this afternoon), but my sense is the majority of the robots present fit into one of two categories: robot vacuums or greeter robots. The two Xiaomi robots -- CyberOne and CyberDog -- may well have been the most prominent of the show, and neither were especially inspiring. It was fun finally seeing the Cyber One in person after writing about it seven months ago. The humanoid robot's stilted locomotion screamed "research prototype" in the first demo, and I'm plenty wary about phone makers getting "serious" about robotics. There was no demo in the booth this year, rendering it more of an expensive mechanical mannequin. CyberDog was moving, at least.
Factuality Enhanced Language Models for Open-Ended Text Generation
Lee, Nayeon, Ping, Wei, Xu, Peng, Patwary, Mostofa, Fung, Pascale, Shoeybi, Mohammad, Catanzaro, Bryan
Pretrained language models (LMs) are susceptible to generate text with nonfactual information. In this work, we measure and improve the factual accuracy of large-scale LMs for open-ended text generation. We design the FactualityPrompts test set and metrics to measure the factuality of LM generations. Based on that, we study the factual accuracy of LMs with parameter sizes ranging from 126M to 530B. Interestingly, we find that larger LMs are more factual than smaller ones, although a previous study suggests that larger LMs can be less truthful in terms of misconceptions. In addition, popular sampling algorithms (e.g., top-p) in open-ended text generation can harm the factuality due to the ''uniform randomness'' introduced at every sampling step. We propose the factual-nucleus sampling algorithm that dynamically adapts the randomness to improve the factuality of generation while maintaining quality. Furthermore, we analyze the inefficiencies of the standard training method in learning correct associations between entities from factual text corpus (e.g., Wikipedia). We propose a factuality-enhanced training method that uses TopicPrefix for better awareness of facts and sentence completion as the training objective, which can vastly reduce the factual errors. We release our code and FactualityPrompts benchmark at: https://github.com/nayeon7lee/FactualityPrompt.
A Planning-Based Explainable Collaborative Dialogue System
Cohen, Philip R., Galescu, Lucian
Eva is a multimodal conversational system that helps users to accomplish their domain goals through collaborative dialogue. The system does this by inferring users' intentions and plans to achieve those goals, detects whether obstacles are present, finds plans to overcome them or to achieve higher-level goals, and plans its actions, including speech acts,to help users accomplish those goals. In doing so, the system maintains and reasons with its own beliefs, goals and intentions, and explicitly reasons about those of its user. Belief reasoning is accomplished with a modal Horn-clause meta-interpreter. The planning and reasoning subsystems obey the principles of persistent goals and intentions, including the formation and decomposition of intentions to perform complex actions, as well as the conditions under which they can be given up. In virtue of its planning process, the system treats its speech acts just like its other actions -- physical acts affect physical states, digital acts affect digital states, and speech acts affect mental and social states. This general approach enables Eva to plan a variety of speech acts including requests, informs, questions, confirmations, recommendations, offers, acceptances, greetings, and emotive expressions. Each of these has a formally specified semantics which is used during the planning and reasoning processes. Because it can keep track of different users' mental states, it can engage in multi-party dialogues. Importantly, Eva can explain its utterances because it has created a plan standing behind each of them. Finally, Eva employs multimodal input and output, driving an avatar that can perceive and employ facial and head movements along with emotive speech acts.
Can ChatGPT Understand Too? A Comparative Study on ChatGPT and Fine-tuned BERT
Zhong, Qihuang, Ding, Liang, Liu, Juhua, Du, Bo, Tao, Dacheng
Recently, ChatGPT has attracted great attention, as it can generate fluent and high-quality responses to human inquiries. Several prior studies have shown that ChatGPT attains remarkable generation ability compared with existing models. However, the quantitative analysis of ChatGPT's understanding ability has been given little attention. In this report, we explore the understanding ability of ChatGPT by evaluating it on the most popular GLUE benchmark, and comparing it with 4 representative fine-tuned BERT-style models. We find that: 1) ChatGPT falls short in handling paraphrase and similarity tasks; 2) ChatGPT outperforms all BERT models on inference tasks by a large margin; 3) ChatGPT achieves comparable performance compared with BERT on sentiment analysis and question-answering tasks. Additionally, by combining some advanced prompting strategies, we show that the understanding ability of ChatGPT can be further improved.
Do we need a National Algorithms Safety Board?
In the United States, the National Transportation Safety Board is widely respected for its prompt responses to investigate plane, train, and boat accidents. Its independent reports have done much to promote safety in civil aviation and beyond. Could a National Algorithms Safety Board have a similar impact in increasing safety for algorithmic systems, especially the rapidly proliferating Artificial Intelligence applications based on unpredictable machine learning? Alternatively, could agencies such as the Food & Drug Administration (FDA), Securities and Exchange Commission (SEC), or Federal Communications Commission (FCC) take on the task of increasing safety of algorithmic systems? In addition to federal agencies, could the major accounting firms provide algorithmic audits as they do in auditing financial statements of publicly listed companies?