Goto

Collaborating Authors

 Media


Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations

arXiv.org Artificial Intelligence

Building socialbots that can have deep, engaging open-domain conversations with humans is one of the grand challenges of artificial intelligence (AI). To this end, bots need to be able to leverage world knowledge spanning several domains effectively when conversing with humans who have their own world knowledge. Existing knowledge-grounded conversation datasets are primarily stylized with explicit roles for conversation partners. These datasets also do not explore depth or breadth of topical coverage with transitions in conversations. We introduce Topical-Chat, a knowledge-grounded human-human conversation dataset where the underlying knowledge spans 8 broad topics and conversation partners don't have explicitly defined roles, to help further research in open-domain conversational AI. We also train several state-of-the-art encoder-decoder conversational models on Topical-Chat and perform automated and human evaluation for benchmarking.


LongDanceDiff: Long-term Dance Generation with Conditional Diffusion Model

arXiv.org Artificial Intelligence

Dancing with music is always an essential human art form to express emotion. Due to the high temporal-spacial complexity, long-term 3D realist dance generation synchronized with music is challenging. Existing methods suffer from the freezing problem when generating long-term dances due to error accumulation and training-inference discrepancy. To address this, we design a conditional diffusion model, LongDanceDiff, for this sequence-to-sequence long-term dance generation, addressing the challenges of temporal coherency and spatial constraint. LongDanceDiff contains a transformer-based diffusion model, where the input is a concatenation of music, past motions, and noised future motions. This partial noising strategy leverages the full-attention mechanism and learns the dependencies among music and past motions. To enhance the diversity of generated dance motions and mitigate the freezing problem, we introduce a mutual information minimization objective that regularizes the dependency between past and future motions. We also address common visual quality issues in dance generation, such as foot sliding and unsmooth motion, by incorporating spatial constraints through a Global-Trajectory Modulation (GTM) layer and motion perceptual losses, thereby improving the smoothness and naturalness of motion generation. Extensive experiments demonstrate a significant improvement in our approach over the existing state-of-the-art methods. We plan to release our codes and models soon.


BAN-PL: a Novel Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl web service

arXiv.org Artificial Intelligence

Advances in automated detection of offensive language online, including hate speech and cyberbullying, require improved access to publicly available datasets comprising social media content. In this paper, we introduce BAN-PL, the first open dataset in the Polish language that encompasses texts flagged as harmful and subsequently removed by professional moderators. The dataset encompasses a total of 691,662 pieces of content from a popular social networking service, Wykop, often referred to as the "Polish Reddit", including both posts and comments, and is evenly distributed into two distinct classes: "harmful" and "neutral". We provide a comprehensive description of the data collection and preprocessing procedures, as well as highlight the linguistic specificity of the data. The BAN-PL dataset, along with advanced preprocessing scripts for, i.a., unmasking profanities, will be publicly available.


A Massive Scale Semantic Similarity Dataset of Historical English

arXiv.org Artificial Intelligence

A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively small datasets created in the past decade by human annotators. This study utilizes a novel source, newly digitized articles from off-copyright, local U.S. newspapers, to assemble a massive-scale semantic similarity dataset spanning 70 years from 1920 to 1989 and containing nearly 400M positive semantic similarity pairs. Historically, around half of articles in U.S. local newspapers came from newswires like the Associated Press. While local papers reproduced articles from the newswire, they wrote their own headlines, which form abstractive summaries of the associated articles. We associate articles and their headlines by exploiting document layouts and language understanding. We then use deep neural methods to detect which articles are from the same underlying source, in the presence of substantial noise and abridgement. The headlines of reproduced articles form positive semantic similarity pairs. The resulting publicly available HEADLINES dataset is significantly larger than most existing semantic similarity datasets and covers a much longer span of time. It will facilitate the application of contrastively trained semantic similarity models to a variety of tasks, including the study of semantic change across space and time.


Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing

arXiv.org Artificial Intelligence

Fashion illustration is used by designers to communicate their vision and to bring the design idea from conceptualization to realization, showing how clothes interact with the human body. In this context, computer vision can thus be used to improve the fashion design process. Differently from previous works that mainly focused on the virtual try-on of garments, we propose the task of multimodal-conditioned fashion image editing, guiding the generation of human-centric fashion images by following multimodal prompts, such as text, human body poses, and garment sketches. We tackle this problem by proposing a new architecture based on latent diffusion models, an approach that has not been used before in the fashion domain. Given the lack of existing datasets suitable for the task, we also extend two existing fashion datasets, namely Dress Code and VITON-HD, with multimodal annotations collected in a semi-automatic manner. Experimental results on these new datasets demonstrate the effectiveness of our proposal, both in terms of realism and coherence with the given multimodal inputs. Source code and collected multimodal annotations are publicly available at: https://github.com/aimagelab/multimodal-garment-designer.


Domain Specific Question Answering Over Knowledge Graphs Using Logical Programming and Large Language Models

arXiv.org Artificial Intelligence

Question Answering over Knowledge Graphs We propose an approach that utilizes LLMs to represent (KGQA) poses significant challenges in the field questions within a specific domain, extracting of Natural Language Processing (NLP). As structured their meanings, while employing logical programming knowledge graphs capturing rich semantic techniques for reasoning and knowledge information become prevalent, there is a pressing representation. Our objective is to demonstrate need for intelligent systems that can reason effectively how this integration enables robust and adaptable and provide accurate answers to intricate KGQA systems that can navigate domain-specific questions within specific domains. The primary knowledge graphs and provide accurate answers to focus of KGQA is to bridge the gap between human complex questions. To evaluate the effectiveness language and structured knowledge representations. of our proposed approach, we conduct experiments When presented with a question in natural using the MetaQA dataset (Zhang et al., 2018), language, KGQA systems aim to traverse the a widely adopted benchmark in KGQA research.


Low-Resource Authorship Style Transfer: Can Non-Famous Authors Be Imitated?

arXiv.org Artificial Intelligence

Authorship style transfer involves altering text to match the style of a target author whilst preserving the original meaning. Existing unsupervised approaches like STRAP have largely focused on style transfer to target authors with many examples of their writing style in books, speeches, or other published works. This high-resource training data requirement (often greater than 100,000 words) makes these approaches primarily useful for style transfer to published authors, politicians, or other well-known figures and authorship styles, while style transfer to non-famous authors has not been well-studied. We introduce the \textit{low-resource authorship style transfer} task, a more challenging class of authorship style transfer where only a limited amount of text in the target author's style may exist. In our experiments, we specifically choose source and target authors from Reddit and style transfer their Reddit posts, limiting ourselves to just 16 posts (on average ~500 words) of the target author's style. Style transfer accuracy is typically measured by how often a classifier or human judge will classify an output as written by the target author. Recent authorship representations models excel at authorship identification even with just a few writing samples, making automatic evaluation of this task possible for the first time through evaluation metrics we propose. Our results establish an in-context learning technique we develop as the strongest baseline, though we find current approaches do not yet achieve mastery of this challenging task. We release our data and implementations to encourage further investigation.


Emergent segmentation from participation dynamics and multi-learner retraining

arXiv.org Artificial Intelligence

The choice to participate in a data-driven service, often made on the basis of quality of that service, influences the ability of the service to learn and improve. We study the participation and retraining dynamics that arise when both the learners and sub-populations of users are \emph{risk-reducing}, which cover a broad class of updates including gradient descent, multiplicative weights, etc. Suppose, for example, that individuals choose to spend their time amongst social media platforms proportionally to how well each platform works for them. Each platform also gathers data about its active users, which it uses to update parameters with a gradient step. For this example and for our general class of dynamics, we show that the only asymptotically stable equilibria are segmented, with sub-populations allocated to a single learner. Under mild assumptions, the utilitarian social optimum is a stable equilibrium. In contrast to previous work, which shows that repeated risk minimization can result in representation disparity and high overall loss for a single learner \citep{hashimoto2018fairness,miller2021outside}, we find that repeated myopic updates with multiple learners lead to better outcomes. We illustrate the phenomena via a simulated example initialized from real data.


911 AI operator weeds out non-emergency calls to free up first responders

FOX News

Former Chicago 911 dispatcher Keith Thornton Jr. joined "Fox & Friends First" to discuss how the crime surge is affecting law enforcement and communities nationwide. Understaffed 911 call centers across the country field non-emergency calls about stray animals or noise complaints on top of their workload of answering serious reports of medical emergencies, crimes and even death. Officials in Charleston County, South Carolina, however, are now leveraging artificial intelligence to streamline non-emergency calls in an effort to free up 911 operators to focus on getting first responders to the scene of emergency incidents as quickly as possible. "Our job is to serve the public the best way we can. So, I am not in any way demeaning anyone from the public, but someone who has their favorite cat stuck in a tree, that's an emergency for them as compared to someone's just been shot," Jim Lake, director of the Charleston County Consolidated Emergency Communications Center, told Fox News Digital in a recent phone interview.


Russian air defences down two drones near Moscow, mayor says

Al Jazeera

Russian air defence systems have brought down two combat drones west of the Russian capital, Moscow mayor Sergei Sobyanin said. The drones were downed early on Tuesday over the Moscow region's towns of Krasnogorsk and the settlement of Chastsy, Sobyanin said. One in the Krasnogorsk area, the other in the Chastsy area," he Sobyanin said on the Telegram messaging app, adding that emergency services were responding. The Moscow mayor did not give details on damage or casualties in what is the latest attempted drone raid on the Russian capital. Air traffic at Moscow's Vnukovo, Sheremetyevo and Domodedovo airports was briefly halted, Russia's state news agency TASS reported, quoting an aviation service source as saying. "Glass damage was recorded on several floors" in a multi-storey residential building in Krasnogorsk," the news agency said, without specifying whether it was the result of a drone strike.