Media
Improving Open-Domain Dialogue Evaluation with a Causal Inference Model
Le, Cat P., Dai, Luke, Johnston, Michael, Liu, Yang, Walker, Marilyn, Ghanadan, Reza
Effective evaluation methods remain a significant challenge for research on open-domain conversational dialogue systems. Explicit satisfaction ratings can be elicited from users, but users often do not provide ratings when asked, and those they give can be highly subjective. Post-hoc ratings by experts are an alternative, but these can be both expensive and complex to collect. Here, we explore the creation of automated methods for predicting both expert and user ratings of open-domain dialogues. We compare four different approaches. First, we train a baseline model using an end-to-end transformer to predict ratings directly from the raw dialogue text. The other three methods are variants of a two-stage approach in which we first extract interpretable features at the turn level that capture, among other aspects, user dialogue behaviors indicating contradiction, repetition, disinterest, compliments, or criticism. We project these features to the dialogue level and train a dialogue-level MLP regression model, a dialogue-level LSTM, and a novel causal inference model called counterfactual-LSTM (CF-LSTM) to predict ratings. The proposed CF-LSTM is a sequential model over turn-level features which predicts ratings using multiple regressors depending on hypotheses derived from the turn-level features. As a causal inference model, CF-LSTM aims to learn the underlying causes of a specific event, such as a low rating. We also bin the user ratings and perform classification experiments with all four models. In evaluation experiments on conversational data from the Alexa Prize SocialBot, we show that the CF-LSTM achieves the best performance for predicting dialogue ratings and classification.
Emergence of Maps in the Memories of Blind Navigation Agents
Wijmans, Erik, Savva, Manolis, Essa, Irfan, Lee, Stefan, Morcos, Ari S., Batra, Dhruv
Animal navigation research posits that organisms build and maintain internal spatial representations, or maps, of their environment. We ask if machines -- specifically, artificial intelligence (AI) navigation agents -- also build implicit (or 'mental') maps. A positive answer to this question would (a) explain the surprising phenomenon in recent literature of ostensibly map-free neural-networks achieving strong performance, and (b) strengthen the evidence of mapping as a fundamental mechanism for navigation by intelligent embodied agents, whether they be biological or artificial. Unlike animal navigation, we can judiciously design the agent's perceptual system and control the learning paradigm to nullify alternative navigation mechanisms. Specifically, we train 'blind' agents -- with sensing limited to only egomotion and no other sensing of any kind -- to perform PointGoal navigation ('go to $\Delta$ x, $\Delta$ y') via reinforcement learning. Our agents are composed of navigation-agnostic components (fully-connected and recurrent neural networks), and our experimental setup provides no inductive bias towards mapping. Despite these harsh conditions, we find that blind agents are (1) surprisingly effective navigators in new environments (~95% success); (2) they utilize memory over long horizons (remembering ~1,000 steps of past experience in an episode); (3) this memory enables them to exhibit intelligent behavior (following walls, detecting collisions, taking shortcuts); (4) there is emergence of maps and collision detection neurons in the representations of the environment built by a blind agent as it navigates; and (5) the emergent maps are selective and task dependent (e.g. the agent 'forgets' exploratory detours). Overall, this paper presents no new techniques for the AI audience, but a surprising finding, an insight, and an explanation.
An Comparative Analysis of Different Pitch and Metrical Grid Encoding Methods in the Task of Sequential Music Generation
Li, Yuqiang, Li, Shengchen, Fazekas, George
Pitch and meter are two fundamental music features for symbolic music generation tasks, where researchers usually choose different encoding methods depending on specific goals. However, the advantages and drawbacks of different encoding methods have not been frequently discussed. This paper presents a integrated analysis of the influence of two low-level feature, pitch and meter, on the performance of a token-based sequential music generation model. First, the commonly used MIDI number encoding and a less used class-octave encoding are compared. Second, an dense intra-bar metric grid is imposed to the encoded sequence as auxiliary features. Different complexity and resolutions of the metric grid are compared. For complexity, the single token approach and the multiple token approach are compared; for grid resolution, 0 (ablation), 1 (bar-level), 4 (downbeat-level) 12, (8th-triplet-level) up to 64 (64th-note-grid-level) are compared; for duration resolution, 4, 8, 12 and 16 subdivisions per beat are compared. All different encodings are tested on separately trained Transformer-XL models for a melody generation task. Regarding distribution similarity of several objective evaluation metrics to the test dataset, results suggest that the class-octave encoding significantly outperforms the taken-for-granted MIDI encoding on pitch-related metrics; finer grids and multiple-token grids improve the rhythmic quality, but also suffer from over-fitting at early training stage. Results display a general phenomenon of over-fitting from two aspects, the pitch embedding space and the test loss of the single-token grid encoding. From a practical perspective, we both demonstrate the feasibility and raise the concern of easy over-fitting problem of using smaller networks and lower embedding dimensions on the generation task. The findings can also contribute to futural models in terms of feature engineering.
A deep-learning search for technosignatures of 820 nearby stars
Ma, Peter Xiangyuan, Ng, Cherry, Rizk, Leandro, Croft, Steve, Siemion, Andrew P. V., Brzycki, Bryan, Czech, Daniel, Drew, Jamie, Gajjar, Vishal, Hoang, John, Isaacson, Howard, Lebofsky, Matt, MacMahon, David, de Pater, Imke, Price, Danny C., Sheikh, Sofia Z., Worden, S. Pete
The goal of the Search for Extraterrestrial Intelligence (SETI) is to quantify the prevalence of technological life beyond Earth via their "technosignatures". One theorized technosignature is narrowband Doppler drifting radio signals. The principal challenge in conducting SETI in the radio domain is developing a generalized technique to reject human radio frequency interference (RFI). Here, we present the most comprehensive deep-learning based technosignature search to date, returning 8 promising ETI signals of interest for re-observation as part of the Breakthrough Listen initiative. The search comprises 820 unique targets observed with the Robert C. Byrd Green Bank Telescope, totaling over 480, hr of on-sky data. We implement a novel beta-Convolutional Variational Autoencoder to identify technosignature candidates in a semi-unsupervised manner while keeping the false positive rate manageably low. This new approach presents itself as a leading solution in accelerating SETI and other transient research into the age of data-driven astronomy.
AI Content Platform MetatronAI.com Announces New Features and Services
DOVER, DE, Jan. 24, 2023 (GLOBE NEWSWIRE) -- Metatron Inc (OTC: MRNJ), an AI content platform, is pleased to announce the release of new services for content generation. MetatronAI.com is a generative artificial intelligence service based on cutting-edge language processing that sets a new standard in the industry. With its advanced natural language understanding capabilities and ability to generate human-like text, art, and soon video and music, designed for business and individuals to create cost-effective original content at unprecedented levels of quality and speed. New features being added on a regular basis include amazing human-like AI text generation for ads, blogs, essays, social media posts, emails, business plans, resumes, books and any type of creative copywriting. Royalty-Free AI art generation is also now live at MetatronAI.com, with professional digital editing coming soon.
World's first AI news anchor unveiled in China
China's state news agency Xinhua this week introduced the newest members of its newsroom: AI anchors who will report "tirelessly" all day every day, from anywhere in the country. Chinese viewers were greeted with a digital version of a regular Xinhua news anchor named Qiu Hao. The anchor, wearing a red tie and pin-striped suit, nods his head in emphasis, blinking and raising his eyebrows slightly. "Not only can I accompany you 24 hours a day, 365 days a year. I can be endlessly copied and present at different scenes to bring you the news," he says.
Mitch Albom: ChatGPT is smart, fast and easy -- all the reasons you should be wary
This is how ChatGPT works. You go to your device, you sign up, it prompts you to ask any question in the world, and you do so. Because what spits back, instantly, is the answer to almost anything, in clear, basic language that sounds like someone is talking to you. Which is kind of the idea. ChatGPT is the latest darling from the world of AI, which, depending on your level of fear, stands for artificial intelligence, allegedly innocent, or alien invasion.
Listen: Google's music-writing AI bot that 'could trick exam setters'
Nello Cristianini, a professor of AI at the University of Bath, said AI has been associated with the creation of music as far back as 1980, however MusicLM is the "most advanced" yet. "It is clearly going to be used, and useful, and controversial too," Prof Cristianini said. "This technology is still unexplored, and we have not tested its legal ramifications." "If you're two musicians who are tasked with composing a piece, and they both happen to stumble across the same AI and happen to put in the same form of words, then presumably they're going to come up with the same product," he said. "I just don't know how you'd unpick AI versus AI?"
Bible by Bot - Sponsored Content
In the last few months, a new AI called ChatGPT has emerged and is already upending education at all levels. How will ChatGPT impact Jewish education and Jewish learning? Identity/Crisis guest host David Zvi Kalman, Director of New Media and Scholar in Residence speaks with Sara Wolkenfeld, Rabbinic Fellow of the David Hartman Center and Chief Learning Officer at Sefaria about what these technologies mean for Jewish learning, how we think about the sacredness of texts, and where we go from here. Identity/Crisis is a weekly podcast from the Shalom Hartman Institute about news and ideas. Subscribe to Identity/Crisis on Apple Podcasts, Spotify or wherever you receive your podcasts.
Reimagining the Communications Industry with ChatGPT -- Sochin Limited
Everybody seems to be enthralled with ChatGPT and the technology is indeed impressive. The chatbot is passing all kinds of graduate level exams, and educational institutions are rushing to devise policies to prevent students from cheating on assessments using the program. ChatGPT can even write scientific abstracts and research papers that fool scientists into thinking that they are real reports. Some users, however, have reported that the chatbot cited non-existent academic journal articles as source material. Amidst all this buzz, we recently conducted our own simple and non-iterative queries to test how the program fares with written texts, which are bread-and-butter products for the communications and PR industry. We noticed that prompts generated texts that included a few paragraphs at most and were short on e.g., persuasive arguments or supporting evidence.