Oceania
FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation
Zhang, Chen, D'Haro, Luis Fernando, Zhang, Qiquan, Friedrichs, Thomas, Li, Haizhou
Recent model-based reference-free metrics for open-domain dialogue evaluation exhibit promising correlations with human judgment. However, they either perform turn-level evaluation or look at a single dialogue quality dimension. One would expect a good evaluation metric to assess multiple quality dimensions at the dialogue level. To this end, we are motivated to propose a multi-dimensional dialogue-level metric, which consists of three sub-metrics with each targeting a specific dimension. The sub-metrics are trained with novel self-supervised objectives and exhibit strong correlations with human judgment for their respective dimensions. Moreover, we explore two approaches to combine the sub-metrics: metric ensemble and multitask learning. Both approaches yield a holistic metric that significantly outperforms individual sub-metrics. Compared to the existing state-of-the-art metric, the combined metrics achieve around 16% relative improvement on average across three high-quality dialogue-level evaluation benchmarks.
Variational Autoencoder with Disentanglement Priors for Low-Resource Task-Specific Natural Language Generation
Li, Zhuang, Qu, Lizhen, Xu, Qiongkai, Wu, Tongtong, Zhan, Tianyang, Haffari, Gholamreza
In this paper, we propose a variational autoencoder with disentanglement priors, VAE-DPRIOR, for task-specific natural language generation with none or a handful of task-specific labeled examples. In order to tackle compositional generalization across tasks, our model performs disentangled representation learning by introducing a conditional prior for the latent content space and another conditional prior for the latent label space. Both types of priors satisfy a novel property called $\epsilon$-disentangled. We show both empirically and theoretically that the novel priors can disentangle representations even without specific regularizations as in the prior work. The content prior enables directly sampling diverse content representations from the content space learned from the seen tasks, and fuse them with the representations of novel tasks for generating semantically diverse texts in the low-resource settings. Our extensive experiments demonstrate the superior performance of our model over competitive baselines in terms of i) data augmentation in continuous zero/few-shot learning, and ii) text style transfer in the few-shot setting.
LEADER: Learning Attention over Driving Behaviors for Planning under Uncertainty
Danesh, Mohamad H., Cai, Panpan, Hsu, David
Uncertainty on human behaviors poses a significant challenge to autonomous driving in crowded urban environments. The partially observable Markov decision processes (POMDPs) offer a principled framework for planning under uncertainty, often leveraging Monte Carlo sampling to achieve online performance for complex tasks. However, sampling also raises safety concerns by potentially missing critical events. To address this, we propose a new algorithm, LEarning Attention over Driving bEhavioRs (LEADER), that learns to attend to critical human behaviors during planning. LEADER learns a neural network generator to provide attention over human behaviors in real-time situations. It integrates the attention into a belief-space planner, using importance sampling to bias reasoning towards critical events. To train the algorithm, we let the attention generator and the planner form a min-max game. By solving the min-max game, LEADER learns to perform risk-aware planning without human labeling.
Learning Invariant Representation and Risk Minimized for Unsupervised Accent Domain Adaptation
Zhao, Chendong, Wang, Jianzong, Qu, Xiaoyang, Wang, Haoqian, Xiao, Jing
Unsupervised representation learning for speech audios attained impressive performances for speech recognition tasks, particularly when annotated speech is limited. However, the unsupervised paradigm needs to be carefully designed and little is known about what properties these representations acquire. There is no guarantee that the model learns meaningful representations for valuable information for recognition. Moreover, the adaptation ability of the learned representations to other domains still needs to be estimated. In this work, we explore learning domain-invariant representations via a direct mapping of speech representations to their corresponding high-level linguistic informations. Results prove that the learned latents not only capture the articulatory feature of each phoneme but also enhance the adaptation ability, outperforming the baseline largely on accented benchmarks.
How Far are We from Robust Long Abstractive Summarization?
Koh, Huan Yee, Ju, Jiaxin, Zhang, He, Liu, Ming, Pan, Shirui
Abstractive summarization has made tremendous progress in recent years. In this work, we perform fine-grained human annotations to evaluate long document abstractive summarization systems (i.e., models and metrics) with the aim of implementing them to generate reliable summaries. For long document abstractive models, we show that the constant strive for state-of-the-art ROUGE results can lead us to generate more relevant summaries but not factual ones. For long document evaluation metrics, human evaluation results show that ROUGE remains the best at evaluating the relevancy of a summary. It also reveals important limitations of factuality metrics in detecting different types of factual errors and the reasons behind the effectiveness of BARTScore. We then suggest promising directions in the endeavor of developing factual consistency metrics. Finally, we release our annotated long document dataset with the hope that it can contribute to the development of metrics across a broader range of summarization settings.
Diverse Parallel Data Synthesis for Cross-Database Adaptation of Text-to-SQL Parsers
Awasthi, Abhijeet, Sathe, Ashutosh, Sarawagi, Sunita
Text-to-SQL parsers typically struggle with databases unseen during the train time. Adapting parsers to new databases is a challenging problem due to the lack of natural language queries in the new schemas. We present ReFill, a framework for synthesizing high-quality and textually diverse parallel datasets for adapting a Text-to-SQL parser to a target schema. ReFill learns to retrieve-and-edit text queries from the existing schemas and transfers them to the target schema. We show that retrieving diverse existing text, masking their schema-specific tokens, and refilling with tokens relevant to the target schema, leads to significantly more diverse text queries than achievable by standard SQL-to-Text generation methods. Through experiments spanning multiple databases, we demonstrate that fine-tuning parsers on datasets synthesized using ReFill consistently outperforms the prior data-augmentation methods.
A Critical Reflection and Forward Perspective on Empathy and Natural Language Processing
Lahnala, Allison, Welch, Charles, Jurgens, David, Flek, Lucie
We review the state of research on empathy in natural language processing and identify the following issues: (1) empathy definitions are absent or abstract, which (2) leads to low construct validity and reproducibility. Moreover, (3) emotional empathy is overemphasized, skewing our focus to a narrow subset of simplified tasks. We believe these issues hinder research progress and argue that current directions will benefit from a clear conceptualization that includes operationalizing cognitive empathy components. Our main objectives are to provide insight and guidance on empathy conceptualization for NLP research objectives and to encourage researchers to pursue the overlooked opportunities in this area, highly relevant, e.g., for clinical and educational sectors.
Design a Sustainable Micro-mobility Future: Trends and Challenges in the United States and European Union Using Natural Language Processing Techniques
Avetisyan, Lilit, Zhang, Chengxin, Bai, Sue, Pari, Ehsan Moradi, Feng, Fred, Bao, Shan, Zhou, Feng
ABSTRACT Micro-mobility is promising to contribute to sustainable cities in the future with its efficiency and low cost. To better design such a sustainable future, it is necessary to understand the trends and challenges. Thus, we examined people's opinions on micro-mobility in the US and the EU using Tweets. We used topic modeling based on advanced natural language processing techniques and categorized the data into seven topics: promotion and service, mobility, technical features, acceptance, recreation, infrastructure and regulations. Furthermore, using sentiment analysis, we investigated people's positive and negative attitudes towards specific aspects of these topics and compared the patterns of the trends and challenges in the US and the EU. We found that 1) promotion and service included the majority of Twitter discussions in the both regions, 2) the EU had more positive opinions than the US, 3) micro-mobility devices were more widely used for utilitarian mobility and recreational purposes in the EU than in the US, and 4) compared to the EU, people in the US had many more concerns related to infrastructure and regulation issues. These findings help us understand the trends and challenges and prioritize different aspects in micro-mobility to improve their safety and experience across the two areas for designing a more sustainable micro-mobility future. INTRODUCTION The growth of transportation has raised the need for compact, flexible, and more sustainable forms of transportation. Recent developments in the micro-mobility industry show that these devices might address this issue and offer people safer and cheaper trips with reduced travel time. According to the Society of Automotive Engineers (SAE) definition (Society of Automotive Engineers, 2019), micro-mobility refers to a range of small, less than 500 pounds (227 kg) lightweight, fully motorized or motor-assisted devices operating at a speed below 30 mph (48 km/h) and ideal for trips up to 10 km. Typical examples include e-bikes, e-scooters, e-unicycles and e-skateboards, and some of them are widely used as personal or shared transportation devices (Price, Blackshear, Blount Jr, & Sandt, 2021). The global micro-mobility market has been increasing over the years. According to the NACTO (National Association of City Transportation Officials, 2020), 136 million trips were generated by shared micro-mobility in 2019 in the U.S., which was 60% more than 2018. Thus, micro-mobility devices can be well integrated into the overall urban design process of smart and sustainable transportation in the near future. With the sustainable design and development goal, we should not only consider technical challenges and requirements (e.g., battery and material), but also complement and constrain the design and development process by social, infrastructural, and political schemes for a sustainable future (Jiao, Luo, Malmqvist, Johan, & Summers, 2022).
On the Global Convergence Rates of Decentralized Softmax Gradient Play in Markov Potential Games
Zhang, Runyu, Mei, Jincheng, Dai, Bo, Schuurmans, Dale, Li, Na
Softmax policy gradient is a popular algorithm for policy optimization in single-agent reinforcement learning, particularly since projection is not needed for each gradient update. However, in multi-agent systems, the lack of central coordination introduces significant additional difficulties in the convergence analysis. Even for a stochastic game with identical interest, there can be multiple Nash Equilibria (NEs), which disables proof techniques that rely on the existence of a unique global optimum. Moreover, the softmax parameterization introduces non-NE policies with zero gradient, making it difficult for gradient-based algorithms in seeking NEs. In this paper, we study the finite time convergence of decentralized softmax gradient play in a special form of game, Markov Potential Games (MPGs), which includes the identical interest game as a special case. We investigate both gradient play and natural gradient play, with and without $\log$-barrier regularization. The established convergence rates for the unregularized cases contain a trajectory-dependent constant that can be arbitrarily large, whereas the $\log$-barrier regularization overcomes this drawback, with the cost of slightly worse dependence on other factors such as the action set size. An empirical study on an identical interest matrix game confirms the theoretical findings.
Using robots to study icebergs
KINGSTON, R.I. – Oct. 27, 2022 – Icebergs originating from Greenland not only impact the environment by affecting sea-level rise and local ecosystems, but they can also damage offshore equipment and disrupt marine transportation. Researchers Mingxi Zhou and Chris Roman from the University of Rhode Island, and Zhuoyuan Song from the University of Hawaiʻi at Mānoa, received a $1.5 million award from the National Science Foundation to develop multi-robot systems for iceberg studies. The research team will develop multiple robotic platforms, including an autonomous surface vehicle, an autonomous underwater vehicle, and several underwater profiling floats, to map the shape of icebergs and assess their surrounding water. They will record the shape, size and drifting speed of the icebergs, and the properties of the surrounding water, creating a unique dataset to better understand iceberg melting and drifting processes. Zhou is the lead researcher on the project.