Goto

Collaborating Authors

 Government


COMBO: A Complete Benchmark for Open KG Canonicalization

arXiv.org Artificial Intelligence

Open knowledge graph (KG) consists of (subject, relation, object) triples extracted from millions of raw text. The subject and object noun phrases and the relation in open KG have severe redundancy and ambiguity and need to be canonicalized. Existing datasets for open KG canonicalization only provide gold entity-level canonicalization for noun phrases. In this paper, we present COMBO, a Complete Benchmark for Open KG canonicalization. Compared with existing datasets, we additionally provide gold canonicalization for relation phrases, gold ontology-level canonicalization for noun phrases, as well as source sentences from which triples are extracted. We also propose metrics for evaluating each type of canonicalization. On the COMBO dataset, we empirically compare previously proposed canonicalization methods as well as a few simple baseline methods based on pretrained language models. We find that properly encoding the phrases in a triple using pretrained language models results in better relation canonicalization and ontology-level canonicalization of the noun phrase. We release our dataset, baselines, and evaluation scripts at https://github.com/jeffchy/COMBO/tree/main.


Fast Linear Model Trees by PILOT

arXiv.org Artificial Intelligence

Linear model trees are regression trees that incorporate linear models in the leaf nodes. This preserves the intuitive interpretation of decision trees and at the same time enables them to better capture linear relationships, which is hard for standard decision trees. But most existing methods for fitting linear model trees are time consuming and therefore not scalable to large data sets. In addition, they are more prone to overfitting and extrapolation issues than standard regression trees. In this paper we introduce PILOT, a new algorithm for linear model trees that is fast, regularized, stable and interpretable. PILOT trains in a greedy fashion like classic regression trees, but incorporates an $L^2$ boosting approach and a model selection rule for fitting linear models in the nodes. The abbreviation PILOT stands for $PI$ecewise $L$inear $O$rganic $T$ree, where `organic' refers to the fact that no pruning is carried out. PILOT has the same low time and space complexity as CART without its pruning. An empirical study indicates that PILOT tends to outperform standard decision trees and other linear model trees on a variety of data sets. Moreover, we prove its consistency in an additive model setting under weak assumptions. When the data is generated by a linear model, the convergence rate is polynomial.


DiversiTree: A New Method to Efficiently Compute Diverse Sets of Near-Optimal Solutions to Mixed-Integer Optimization Problems

arXiv.org Artificial Intelligence

While most methods for solving mixed-integer optimization problems compute a single optimal solution, a diverse set of near-optimal solutions can often lead to improved outcomes. We present a new method for finding a set of diverse solutions by emphasizing diversity within the search for near-optimal solutions. Specifically, within a branch-and-bound framework, we investigated parameterized node selection rules that explicitly consider diversity. Our results indicate that our approach significantly increases the diversity of the final solution set. When compared with two existing methods, our method runs with similar runtime as regular node selection methods and gives a diversity improvement between 12% and 190%. In contrast, popular node selection rules, such as best-first search, in some instances performed worse than state-of-the-art methods by more than 35% and gave an improvement of no more than 130%. Further, we find that our method is most effective when diversity in node selection is continuously emphasized after reaching a minimal depth in the tree and when the solution set has grown sufficiently large. Our method can be easily incorporated into integer programming solvers and has the potential to significantly increase the diversity of solution sets.


A Survey on XAI for Beyond 5G Security: Technical Aspects, Use Cases, Challenges and Research Directions

arXiv.org Artificial Intelligence

With the advent of 5G commercialization, the need for more reliable, faster, and intelligent telecommunication systems are envisaged for the next generation beyond 5G (B5G) radio access technologies. Artificial Intelligence (AI) and Machine Learning (ML) are not just immensely popular in the service layer applications but also have been proposed as essential enablers in many aspects of B5G networks, from IoT devices and edge computing to cloud-based infrastructures. However, existing B5G ML-security surveys tend to place more emphasis on AI/ML model performance and accuracy than on the models' accountability and trustworthiness. In contrast, this paper explores the potential of Explainable AI (XAI) methods, which would allow B5G stakeholders to inspect intelligent black-box systems used to secure B5G networks. The goal of using XAI in the security domain of B5G is to allow the decision-making processes of the ML-based security systems to be transparent and comprehensible to B5G stakeholders making the systems accountable for automated actions. In every facet of the forthcoming B5G era, including B5G technologies such as RAN, zero-touch network management, E2E slicing, this survey emphasizes the role of XAI in them and the use cases that the general users would ultimately enjoy. Furthermore, we presented the lessons learned from recent efforts and future research directions on top of the currently conducted projects involving XAI.


Global Performance Disparities Between English-Language Accents in Automatic Speech Recognition

arXiv.org Artificial Intelligence

However, many users are familiar with the frustrating experience of repeatedly not being understood by their voice assistant [16], so much so that frustration with ASR has become a culturally-shared source of comedy [4, 32]. Bias auditing of ASR services has quantified these experiences. English language ASR has higher error rates: for Black Americans compared to white Americans [24, 45], for stigmatised British accents compared to favored British accents [28], for Scottish speakers compared to speakers from California and New Zealand [44], for speakers whose first language is a tone language compared to those whose first language is not [2], for speakers with Indian accents compared to speakers who with "American" accents [31], for speakers whose first language is English compared to those for whom it is not [28]. It should go without saying, but everyone has an accent - there is no "unaccented" version of English [26]. Due to colonization and globalization, different Englishes are spoken around the world. While some English accents may be favored by those with class, race, and national origin privilege [28], there is no technical barrier to building an ASR system which works well on any particular accent. So we are left with the question, why does ASR performance vary as it does as a function of the global English accent spoken?


ChatGPT Has Been Sucked Into India's Culture Wars

WIRED

A tweet pinned to the top of Hegde's feed in honor of Modi's birthday calls him "the leader who brought back India's lost glory." On January 7, the account tweeted a screenshot from ChatGPT to its more than 185,000 followers; the tweet appeared to show the AI-powered chatbot making a joke about the Hindu deity Krishna. ChatGPT uses large language models to provide detailed answers to text prompts, responding to questions about everything from legal problems to song lyrics. But on questions of faith, it's mostly trained to be circumspect, responding "I'm sorry, but I'm not programmed to make jokes about any religion or deity," when prompted to quip about Jesus Christ or Mohammed. That limitation appears not to include Hindu religious figures.


How Deepfake Videos Are Used to Spread Disinformation - The New York Times

#artificialintelligence

Their voices were stilted and failed to sync with the movement of their mouths. Their faces had a pixelated, video-game quality and their hair appeared unnaturally plastered to the head. The captions were filled with grammatical mistakes. The two broadcasters, purportedly anchors for a news outlet called Wolf News, are not real people. They are computer-generated avatars created by artificial intelligence software.


The Humanoid Robot NASA Is Helping Build - CNET

#artificialintelligence

We've seen impressive developments in humanoid robots over the last few years. Elon Musk and Tesla introduced the Optimus robot last year, and every few months Boston Dynamics teaches its Atlas robot a few new tricks. Next month at South by Southwest, a Texas-based startup will reveal to a small group its take on a general-purpose robot. Apptronik calls its newest robot Apollo, in part because it partnered with NASA on commercializing the robot. Though there aren't plans to send Apollo to space, the space agency wants to encourage the development of humanoid robots that could one day lead to a robotic space-explorer.


Data Analyst, Execution, CTR at Standard Bank Group - Johannesburg, South Africa

#artificialintelligence

To conduct regulatory monitoring within Consumer and High Net Worth on a specific set of regulatory requirements (e.g., PEPS, Sanctions, EDD, FIC Amendment Bill, CTR, Waterfall (KYC), AML Training, Quality Assurance, etc.) as prescribed by the Regulatory Monitoring framework and drives first level of defence remediation of breaches. To provide insights on the state of regulatory adherence within allocated portfolio and prepare appropriate reports as input into overall regulatory reporting.


The Age of AI Hacking Is Closer Than You Think

WIRED

How realistic is a future of AI hacking? If you buy something using links in our stories, we may earn a commission. This helps support our journalism. Its feasibility depends on the specific system being modeled and hacked. For an AI to even begin optimizing a solution, let alone develop a completely novel one, all of the rules of the environment must be formalized in a way the computer can understand.