Goto

Collaborating Authors

 Government


Cross-strait Variations on Two Near-synonymous Loanwords xie2shang1 and tan2pan4: A Corpus-based Comparative Study

arXiv.org Artificial Intelligence

This study attempts to investigate cross-strait variations on two typical synonymous loanwords in Chinese, i.e. xie2shang1 and tan2pan4, drawn on MARVS theory. Through a comparative analysis, the study found some distributional, eventual, and contextual similarities and differences across Taiwan and Mainland Mandarin. Compared with the underused tan2pan4, xie2shang1 is significantly overused in Taiwan Mandarin and vice versa in Mainland Mandarin. Additionally, though both words can refer to an inchoative process in Mainland and Taiwan Mandarin, the starting point for xie2shang1 in Mainland Mandarin is somewhat blurring compared with the usage in Taiwan Mandarin. Further on, in Taiwan Mandarin, tan2pan4 can be used in economic and diplomatic contexts, while xie2shang1 is used almost exclusively in political contexts. In Mainland Mandarin, however, the two words can be used in a hybrid manner within political contexts; moreover, tan2pan4 is prominently used in diplomatic contexts with less reference to economic activities, while xie2sahng1 can be found in both political and legal contexts, emphasizing a role of mediation.


Understanding and Improving Zero-shot Multi-hop Reasoning in Generative Question Answering

arXiv.org Artificial Intelligence

Generative question answering (QA) models generate answers to questions either solely based on the parameters of the model (the closed-book setting) or additionally retrieving relevant evidence (the open-book setting). Generative QA models can answer some relatively complex questions, but the mechanism through which they do so is still poorly understood. We perform several studies aimed at better understanding the multi-hop reasoning capabilities of generative QA models. First, we decompose multi-hop questions into multiple corresponding single-hop questions, and find marked inconsistency in QA models' answers on these pairs of ostensibly identical question chains. Second, we find that models lack zero-shot multi-hop reasoning ability: when trained only on single-hop questions, models generalize poorly to multi-hop questions. Finally, we demonstrate that it is possible to improve models' zero-shot multi-hop reasoning capacity through two methods that approximate real multi-hop natural language (NL) questions by training on either concatenation of single-hop questions or logical forms (SPARQL). In sum, these results demonstrate that multi-hop reasoning does not emerge naturally in generative QA models, but can be encouraged by advances in training or modeling techniques.


Forecasting Future World Events with Neural Networks

arXiv.org Artificial Intelligence

Forecasting future world events is a challenging but valuable task. Forecasts of climate, geopolitical conflict, pandemics and economic indicators help shape policy and decision making. In these domains, the judgment of expert humans contributes to the best forecasts. Given advances in language modeling, can these forecasts be automated? To this end, we introduce Autocast, a dataset containing thousands of forecasting questions and an accompanying news corpus. Questions are taken from forecasting tournaments, ensuring high quality, real-world importance, and diversity. The news corpus is organized by date, allowing us to precisely simulate the conditions under which humans made past forecasts (avoiding leakage from the future). Motivated by the difficulty of forecasting numbers across orders of magnitude (e.g. global cases of COVID-19 in 2022), we also curate IntervalQA, a dataset of numerical questions and metrics for calibration. We test language models on our forecasting task and find that performance is far below a human expert baseline. However, performance improves with increased model size and incorporation of relevant information from the news corpus. In sum, Autocast poses a novel challenge for large language models and improved performance could bring large practical benefits.


Automated Clinical Coding: What, Why, and Where We Are?

arXiv.org Artificial Intelligence

Clinical coding is the task of transforming medical information in a patient's health records into structured codes so that they can be used for statistical analysis. This is a cognitive and time-consuming task that follows a standard process in order to achieve a high level of consistency. Clinical coding could potentially be supported by an automated system to improve the efficiency and accuracy of the process. We introduce the idea of automated clinical coding and summarise its challenges from the perspective of Artificial Intelligence (AI) and Natural Language Processing (NLP), based on the literature, our project experience over the past two and half years (late 2019 - early 2022), and discussions with clinical coding experts in Scotland and the UK. Our research reveals the gaps between the current deep learning-based approach applied to clinical coding and the need for explainability and consistency in real-world practice. Knowledge-based methods that represent and reason the standard, explainable process of a task may need to be incorporated into deep learning-based methods for clinical coding. Automated clinical coding is a promising task for AI, despite the technical and organisational challenges. Coders are needed to be involved in the development process. There is much to achieve to develop and deploy an AI-based automated system to support coding in the next five years and beyond.


Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation

arXiv.org Artificial Intelligence

While large-scale neural language models, such as GPT2 and BART, have achieved impressive results on various text generation tasks, they tend to get stuck in undesirable sentence-level loops with maximization-based decoding algorithms (\textit{e.g.}, greedy search). This phenomenon is counter-intuitive since there are few consecutive sentence-level repetitions in human corpora (e.g., 0.02\% in Wikitext-103). To investigate the underlying reasons for generating consecutive sentence-level repetitions, we study the relationship between the probabilities of the repetitive tokens and their previous repetitions in the context. Through our quantitative experiments, we find that 1) Language models have a preference to repeat the previous sentence; 2) The sentence-level repetitions have a \textit{self-reinforcement effect}: the more times a sentence is repeated in the context, the higher the probability of continuing to generate that sentence; 3) The sentences with higher initial probabilities usually have a stronger self-reinforcement effect. Motivated by our findings, we propose a simple and effective training method \textbf{DITTO} (Pseu\underline{D}o-Repet\underline{IT}ion Penaliza\underline{T}i\underline{O}n), where the model learns to penalize probabilities of sentence-level repetitions from pseudo repetitive data. Although our method is motivated by mitigating repetitions, experiments show that DITTO not only mitigates the repetition issue without sacrificing perplexity, but also achieves better generation quality. Extensive experiments on open-ended text generation (Wikitext-103) and text summarization (CNN/DailyMail) demonstrate the generality and effectiveness of our method.


Where do Models go Wrong? Parameter-Space Saliency Maps for Explainability

arXiv.org Artificial Intelligence

Conventional saliency maps highlight input features to which neural network predictions are highly sensitive. We take a different approach to saliency, in which we identify and analyze the network parameters, rather than inputs, which are responsible for erroneous decisions. We first verify that identified salient parameters are indeed responsible for misclassification by showing that turning these parameters off improves predictions on the associated samples, more than turning off the same number of random or least salient parameters. We further validate the link between salient parameters and network misclassification errors by observing that fine-tuning a small number of the most salient parameters on a single sample results in error correction on other samples which were misclassified for similar reasons - nearest neighbors in the saliency space. After validating our parameter-space saliency maps, we demonstrate that samples which cause similar parameters to malfunction are semantically similar. Further, we introduce an input-space saliency counterpart which reveals how image features cause specific network components to malfunction.


US officials meet with Taliban in person for first time since drone strike killed Al Qaeda chief in Kabul

FOX News

Former Army Ranger and Save Our Allies co-founder Tim Kennedy discusses the trauma experienced by military veterans following the evacuation of Afghanistan on'Fox News Live.' Top U.S. officials held their first in person meeting with the Taliban since a U.S. military strike killed the leader of Al Qaeda in Afghanistan in July. The Biden administration sent CIA deputy director David Cohen to the Qatari capital of Doha on Saturday to meet with a Taliban delegation led by Abdul Haq Wasiq, the Taliban's head of intelligence, Fox News has confirmed. The meeting marks the first time the two sides have met in person since a U.S. drone strike this summer killed Al Qaeda leader Ayman al-Zawahri in the Taliban controlled Afghanistan capital of Kabul raising questions about the terror group's presence in the country. The Taliban claimed it was unaware that the Al-Qaeda chief was in the country and called the drone strike a "clear violation" of the Doha agreement struck with former President Donald Trump in 2020. AFGHANISTAN ONE YEAR LATER: HOW DAILY LIFE IN THE WAR-TORN COUNTRY HAS CHANGED SINCE THE TALIBAN'S TAKEOVER Taliban fighters escort women march in support of the Taliban government outside Kabul University, Afghanistan.


Cybersecurity Will Account for Nearly One-Quarter of AI Software Market Through 2025

#artificialintelligence

By 2025, the artificial intelligence (AI) software market will expand from 2021's $33 billion to $64 billion, according to a new report. And cybersecurity is the fastest-growing category of AI spend, experiencing a rise in spending of 22.3% compound annual growth rate (CAGR). "Cybersecurity is the fastest AI software growth category, with a focus on the real-time monitoring of and response to attacks," the report states. The next two categories, customer and human capital management (22%) and process optimization, knowledge, and data intelligence (18.3%), also have cybersecurity elements, so the impact on security tool makers could be even more significant. This comports with the emphasis companies have placed on their AI-enhanced software and services.


October 2022 Issue: Why NASA is sending a surgical robot to the ISS - The Robot Report

#artificialintelligence

Teams must embrace a multileveled approach to product development and employ "deliberate innovation" that flows from the ideation stages all the way to commercialization.


Tech firms say laws to protect us from bad AI will limit 'innovation'. Well, good John Naughton

The Guardian

Way back in May 2014, the European court of justice issued a landmark ruling that European citizens had the right to petition search engines to remove search results that linked to material that had been posted lawfully on third-party websites. This was popularly but misleadingly described as the "right to be forgotten"; it was really a right to have certain published material about the complainant delisted by search engines, of which Google was by far the most dominant. Or, to put it crudely, a right not to be found by Google. On the morning the ruling was released, I had a phone call from a relatively senior Google employee whom I happened to know. It was clear from his call that the company had been ambushed by the ruling – its expensive legal team had plainly not expected it. But it was also clear that his US bosses were incensed by the effrontery of a mere European institution in issuing such a verdict.