Goto

Collaborating Authors

 Government


A Critical Assessment of Interpretable and Explainable Machine Learning for Intrusion Detection

arXiv.org Artificial Intelligence

There has been a large number of studies in interpretable and explainable ML for cybersecurity, in particular, for intrusion detection. Many of these studies have significant amount of overlapping and repeated evaluations and analysis. At the same time, these studies overlook crucial model, data, learning process, and utility related issues and many times completely disregard them. These issues include the use of overly complex and opaque ML models, unaccounted data imbalances and correlated features, inconsistent influential features across different explanation methods, the inconsistencies stemming from the constituents of a learning process, and the implausible utility of explanations. In this work, we empirically demonstrate these issues, analyze them and propose practical solutions in the context of feature-based model explanations. Specifically, we advise avoiding complex opaque models such as Deep Neural Networks and instead using interpretable ML models such as Decision Trees as the available intrusion datasets are not difficult for such interpretable models to classify successfully. Then, we bring attention to the binary classification metrics such as Matthews Correlation Coefficient (which are well-suited for imbalanced datasets. Moreover, we find that feature-based model explanations are most often inconsistent across different settings. In this respect, to further gauge the extent of inconsistencies, we introduce the notion of cross explanations which corroborates that the features that are determined to be impactful by one explanation method most often differ from those by another method. Furthermore, we show that strongly correlated data features and the constituents of a learning process, such as hyper-parameters and the optimization routine, become yet another source of inconsistent explanations. Finally, we discuss the utility of feature-based explanations.


Uncertainty-Guided Optimization on Large Language Model Search Trees

arXiv.org Artificial Intelligence

Beam search is a standard tree search algorithm when it comes to finding sequences of maximum likelihood, for example, in the decoding processes of large language models. However, it is myopic since it does not take the whole path from the root to a leaf into account. Moreover, it is agnostic to prior knowledge available about the process: For example, it does not consider that the objective being maximized is a likelihood and thereby has specific properties, like being bound in the unit interval. Taking a probabilistic approach, we define a prior belief over the LLMs' transition probabilities and obtain a posterior belief over the most promising paths in each iteration. These beliefs are helpful to define a non-myopic Bayesian-optimization-like acquisition function that allows for a more data-efficient exploration scheme than standard beam search. We discuss how to select the prior and demonstrate in on- and off-model experiments with recent large language models, including Llama-2-7b, that our method achieves higher efficiency than beam search: Our method achieves the same or a higher likelihood while expanding fewer nodes than beam search.


L+M-24: Building a Dataset for Language + Molecules @ ACL 2024

arXiv.org Artificial Intelligence

Language-molecule models have emerged as an exciting direction for molecular discovery and understanding. However, training these models is challenging due to the scarcity of molecule-language pair datasets. At this point, datasets have been released which are 1) small and scraped from existing databases, 2) large but noisy and constructed by performing entity linking on the scientific literature, and 3) built by converting property prediction datasets to natural language using templates. In this document, we detail the $\textit{L+M-24}$ dataset, which has been created for the Language + Molecules Workshop shared task at ACL 2024. In particular, $\textit{L+M-24}$ is designed to focus on three key benefits of natural language in molecule design: compositionality, functionality, and abstraction.


FAIR: Filtering of Automatically Induced Rules

arXiv.org Artificial Intelligence

The availability of large annotated data can be a critical bottleneck in training machine learning algorithms successfully, especially when applied to diverse domains. Weak supervision offers a promising alternative by accelerating the creation of labeled training data using domain-specific rules. However, it requires users to write a diverse set of high-quality rules to assign labels to the unlabeled data. Automatic Rule Induction (ARI) approaches circumvent this problem by automatically creating rules from features on a small labeled set and filtering a final set of rules from them. In the ARI approach, the crucial step is to filter out a set of a high-quality useful subset of rules from the large set of automatically created rules. In this paper, we propose an algorithm (Filtering of Automatically Induced Rules) to filter rules from a large number of automatically induced rules using submodular objective functions that account for the collective precision, coverage, and conflicts of the rule set. We experiment with three ARI approaches and five text classification datasets to validate the superior performance of our algorithm with respect to several semi-supervised label aggregation approaches. Further, we show that achieves statistically significant results in comparison to existing rule-filtering approaches.


20 tech tricks to make life better, safer or easier

FOX News

Alex Galvagni, CEO of Age of Learning and a former artificial intelligence researcher with NASA, says advances in AI now make it possible to deliver to children "a personalized and supportive" experience in education. Our everyday devices get new updates and features all the time. It's tough to keep up, but that's why you have me. Below you'll find 20 sweet shortcuts -- some new, some hidden gems that have been there all along. Try my free tech newsletter to enter!


Meta Has Been Ordered to Stop Mining Brazilian Personal Data to Train Its AI

TIME - Tech

Brazil's national data protection authority has ordered Meta to halt the use of data originating from the country to train its AI models. Meta's current privacy policy enables the company to use data from its platforms, including Facebook, Instagram, and WhatsApp to train its artificial intelligence models. However, that practice will no longer be permitted in Brazil after its national data protection authority gave the company five days to change its policy on Tuesday. Brazil said the company will need to confirm it has stopped using the data or face a daily non-compliance fine of 50,000 Brazilian Reals (almost 9000), citing "the imminent risk of serious and irreparable or difficult-to-repair damage to the fundamental rights of the affected data subjects." Meta said it was "disappointed" with the Brazilian authority's decision, saying it was a "step backward for innovation."


White House, Biden campaign call Trump's cognitive ability into question

FOX News

Rep. Jake Auchincloss, D-Mass., joins'America's Newsroom' to discuss reports of Democratic concerns about President Biden's 2024 candidacy. The White House and Biden campaign have questioned former President Trump's fitness to serve as questions continue to swirl around President Biden's mental acuity. Asked during a news conference Tuesday if Biden had Alzheimer's or any form of dementia, White House press secretary Karine Jean-Pierre said "no" while hinting that the "same exact question" should be asked of the "other guy," referring to Trump. The answer comes after Biden continues to face widespread skepticism about his ability to win the election and serve another term as president in the wake of a disastrous debate performance last week, resulting in many calling on the president to step aside and let a younger candidate take over at the top of the ticket. The Biden campaign has acknowledged the president's poor performance but pushed back against the idea he would drop out of the race, arguing Biden still has the ability to lead and is the party's best chance at defeating Trump. The campaign has also begun calling Trump's cognitive ability into question, citing times the former president has confused who he was talking about.


These Weird Pics of Donald Trump Have a Much Darker Backstory

Slate

Who is the president of the United States? But one popular artificial intelligence app thinks otherwise. While A.I. models have been known to hallucinate (i.e., make stuff up), they typically don't mess up extremely simple things like name the president. If you ask OpenAI's ChatGPT who the U.S. president is, it'll give you the correct answer, a short bio, and links to the White House website and Biden's Wikipedia article. The same is true for Anthropic's Claude and Meta AI, while Google's Gemini straight-up refuses to answer the question because it's related to politics.


Fox News AI Newsletter: AI exoskeletons assist performance

FOX News

Alex Galvagni, CEO of Age of Learning and a former artificial intelligence researcher with NASA, says advances in AI now make it possible to deliver to children "a personalized and supportive" experience in education. ROBOTIC POWER WEAR: A groundbreaking AI-powered exoskeleton developed by researchers at North Carolina State University and the University of North Carolina at Chapel Hill promises to be a game-changer for individuals with mobility issues. ELECTION SEASON: Google on Monday announced that it will have a mandatory requirement for advertisers to disclose election ads that use digitally altered content in depictions of real or realistic-looking people or events. Victor Miller is running for mayor of Cheyenne as AI bot'VIC' (Fox News Digital) 'AI FOR MAYOR': A Wyoming man who filed for the state capital's mayor's race as an AI bot named "VIC" spoke to Fox News Digital this week about VIC's landmark candidacy and a breaking setback he encountered moments before taping. SAFEGUARD SUMMER SOJOURNS: A new study by online protection company McAfee has identified the top five destinations most frequently targeted by cybercriminals for online booking scams.


Ukrainian maritime attack on Black Sea port Novorossiysk repelled: Russia

Al Jazeera

Russia says it destroyed two Ukrainian sea drones targeting the Black Sea port of Novorossiysk, a key naval base and oil shipping outlet. The Ministry of Defence in Moscow said on Wednesday that Russian forces had destroyed the naval drones as they advanced on the port in an overnight attack. Ukraine has reported success in targeting Russian ships and infrastructure in the Black Sea over recent months. "Two unmanned boats travelling in the direction of Novorossiysk were destroyed in the waters of the Black Sea," the ministry said in a post on Telegram. The attack caused no damage or shipping disruptions, the local city administration reported, according to Russian state news agencies.