Goto

Collaborating Authors

 Government


Metastatic Breast Cancer Prognostication Through Multimodal Integration of Dimensionality Reduction Algorithms and Classification Algorithms

arXiv.org Artificial Intelligence

Machine learning (ML) is a branch of Artificial Intelligence (AI) where computers analyze data and find patterns in the data. The study focuses on the detection of metastatic cancer using ML. Metastatic cancer is the point where the cancer has spread to other parts of the body and is the cause of approximately 90% of cancer related deaths. Normally, pathologists spend hours each day to manually classify whether tumors are benign or malignant. This tedious task contributes to mislabeling metastasis being over 60% of time and emphasizes the importance to be aware of human error, and other inefficiencies. ML is a good candidate to improve the correct identification of metastatic cancer saving thousands of lives and can also improve the speed and efficiency of the process thereby taking less resources and time. So far, deep learning methodology of AI has been used in the research to detect cancer. This study is a novel approach to determine the potential of using preprocessing algorithms combined with classification algorithms in detecting metastatic cancer. The study used two preprocessing algorithms: principal component analysis (PCA) and the genetic algorithm to reduce the dimensionality of the dataset, and then used three classification algorithms: logistic regression, decision tree classifier, and k-nearest neighbors to detect metastatic cancer in the pathology scans. The highest accuracy of 71.14% was produced by the ML pipeline comprising of PCA, the genetic algorithm, and the k-nearest neighbors algorithm, suggesting that preprocessing and classification algorithms have great potential for detecting metastatic cancer.


LLM4Jobs: Unsupervised occupation extraction and standardization leveraging Large Language Models

arXiv.org Artificial Intelligence

Automated occupation extraction and standardization from free-text job postings and resumes are crucial for applications like job recommendation and labor market policy formation. This paper introduces LLM4Jobs, a novel unsupervised methodology that taps into the capabilities of large language models (LLMs) for occupation coding. LLM4Jobs uniquely harnesses both the natural language understanding and generation capacities of LLMs. Evaluated on rigorous experimentation on synthetic and real-world datasets, we demonstrate that LLM4Jobs consistently surpasses unsupervised state-of-the-art benchmarks, demonstrating its versatility across diverse datasets and granularities. As a side result of our work, we present both synthetic and real-world datasets, which may be instrumental for subsequent research in this domain. Overall, this investigation highlights the promise of contemporary LLMs for the intricate task of occupation extraction and standardization, laying the foundation for a robust and adaptable framework relevant to both research and industrial contexts.


A Survey on Privacy in Graph Neural Networks: Attacks, Preservation, and Applications

arXiv.org Artificial Intelligence

Privacy attack is a popular and well-developed topic in various fields such as social network analysis, healthcare, finance, system, etc. [88], [89], [90]. During recent years, the surge of machine learning has provided powerful tools to solve many practical problems. However, data-driven approaches also threaten users' privacy due to the associated risks of data leakage and inference [85]. Consequently, a substantial amount of work has been devoted to investigate the vulnerabilities of ML models and the risks of privacy leakage [47]. A branch of privacy research is to develop privacy attack models, which has received much attention during the past few years. However, attack models with respect to GNNs have only been explored very recently because GNN techniques are relatively new compared with CNN/transformers in image/natural language processing(NLP) domains, and the irregular graph structure poses unique challenges to transfer existing attack techniques that are well-established in other domains. In this section, we summarize papers that have developed attack models specifically targeting GNNs. Figure 1: Illustrations of the four categories of privacy attack We classify the privacy attack models on GNN into models on graphs: a) Model extraction attacks (MEA); b) four categories (which are visualized in Figure 4): a) model Graph structure reconstruction (GSR); c) Attribute inference extraction attack (MEA), b) graph structure reconstruction attacks (AIA); and d) Membership inference attacks (MIA).


Community Detection Using Revised Medoid-Shift Based on KNN

arXiv.org Artificial Intelligence

Community detection becomes an important problem with the booming of social networks. The Medoid-Shift algorithm preserves the benefits of Mean-Shift and can be applied to problems based on distance matrix, such as community detection. One drawback of the Medoid-Shift algorithm is that there may be no data points within the neighborhood region defined by a distance parameter. To deal with the community detection problem better, a new algorithm called Revised Medoid-Shift (RMS) in this work is thus proposed. During the process of finding the next medoid, the RMS algorithm is based on a neighborhood defined by KNN, while the original Medoid-Shift is based on a neighborhood defined by a distance parameter. Since the neighborhood defined by KNN is more stable than the one defined by the distance parameter in terms of the number of data points within the neighborhood, the RMS algorithm may converge more smoothly. In the RMS method, each of the data points is shifted towards a medoid within the neighborhood defined by KNN. After the iterative process of shifting, each of the data point converges into a cluster center, and the data points converging into the same center are grouped into the same cluster. The RMS algorithm is tested on two kinds of datasets including community datasets with known ground truth partition and community datasets without ground truth partition respectively. The experiment results show sthat the proposed RMS algorithm generally produces betster results than Medoid-Shift and some state-of-the-art together with most classic community detection algorithms on different kinds of community detection datasets.


Deep Kernel Methods Learn Better: From Cards to Process Optimization

arXiv.org Artificial Intelligence

The ability of deep learning methods to perform classification and regression tasks relies heavily on their capacity to uncover manifolds in high-dimensional data spaces and project them into low-dimensional representation spaces. In this study, we investigate the structure and character of the manifolds generated by classical variational autoencoder (VAE) approaches and deep kernel learning (DKL). In the former case, the structure of the latent space is determined by the properties of the input data alone, while in the latter, the latent manifold forms as a result of an active learning process that balances the data distribution and target functionalities. We show that DKL with active learning can produce a more compact and smooth latent space which is more conducive to optimization compared to previously reported methods, such as the VAE. We demonstrate this behavior using a simple cards data set and extend it to the optimization of domain-generated trajectories in physical systems. Our findings suggest that latent manifolds constructed through active learning have a more beneficial structure for optimization problems, especially in feature-rich target-poor scenarios that are common in domain sciences, such as materials synthesis, energy storage, and molecular discovery. The jupyter notebooks that encapsulate the complete analysis accompany the article.


Explainable Deep Learning Methods in Medical Image Classification: A Survey

arXiv.org Artificial Intelligence

The progress made on the last decade in the field of artificial intelligence (AI) has supported a dramatic increase in the accuracy of most computer vision applications. Medical image analysis is one of the applications where the progress made assured human-level accuracy on the classification of different types of medical data (e.g., chest X-rays [90], corneal images [166]). However, and in spite of these advances, automated medical imaging is seldom adopted in clinical practice. According to Zachary Lipton [77], the explanation to this apparent paradox is straightforward, doctors will never trust the decision of an algorithm without understanding its decision process. This fact has raised the need for producing strategies capable of explaining the decision process of AI algorithms, leading subsequently to the creation of a novel research topic named as eXplainable Artificial Intelligence (XAI). According to DARPA [46], XAI aims to "produce more explainable models, while maintaining a high level of learning performance (prediction accuracy); and enable human users to understand, appropriately, trust, and effectively manage the emerging generation of artificially intelligent partners". In spite of its general applicability, XAI is particularly important in high-stake decisions, such as clinical workflow, where the consequences of a wrong decision could lead to human deaths. This is also evidenced by European Union's General Data Protection


PAGER: A Framework for Failure Analysis of Deep Regression Models

arXiv.org Machine Learning

Safe deployment of AI models requires proactive detection of potential prediction failures to prevent costly errors. While failure detection in classification problems has received significant attention, characterizing failure modes in regression tasks is more complicated and less explored. Existing approaches rely on epistemic uncertainties or feature inconsistency with the training distribution to characterize model risk. However, we show that uncertainties are necessary but insufficient to accurately characterize failure, owing to the various sources of error. In this paper, we propose PAGER (Principled Analysis of Generalization Errors in Regressors), a framework to systematically detect and characterize failures in deep regression models. Built upon the recently proposed idea of anchoring in deep models, PAGER unifies both epistemic uncertainties and novel, complementary non-conformity scores to organize samples into different risk regimes, thereby providing a comprehensive analysis of model errors. Additionally, we introduce novel metrics for evaluating failure detectors in regression tasks. We demonstrate the effectiveness of PAGER on synthetic and real-world benchmarks. Our results highlight the capability of PAGER to identify regions of accurate generalization and detect failure cases in out-of-distribution and out-of-support scenarios.


Netanyahu talks to Elon Musk in California about anti-Semitism on X

Al Jazeera

Prime Minister Benjamin Netanyahu is starting a US trip in California to talk about technology and artificial intelligence with billionaire businessman Elon Musk. The Israeli leader posted Monday on Musk's social media platform X, formerly known as Twitter, that he plans to talk with the Tesla CEO "about how we can harness the opportunities and mitigate the risks of AI for the good of civilization." Netanyahu's high-profile visit to the San Francisco Bay Area comes at a time when Musk is facing accusations of tolerating anti-Semitic messages on his social media platform, while Netanyahu is confronting political opposition at home and abroad. Protesters gathered early Monday outside the Fremont, California factory where Tesla makes its cars. The video livestream kicked off shortly before 9:30am with Netanyahu and the Tesla CEO.


Sen. Richard Blumenthal Defends His Controversial Bill Regulating Social Media for Kids

Slate

For a while now, Washington has been wrestling with two big forces shaping technology: social media and artificial intelligence. Who should do it--and how? Currently, Congress is considering a bill that would regulate how social media companies treat minors: the Kids Online Safety Act. Although it has bipartisan support, KOSA is not without controversy. Several critics have called it "government censorship." One group, the Electronic Frontier Foundation, says it is "one of the most dangerous bills in years."


Airport in Iraq's Kurdish region hit by deadly drone attack

Al Jazeera

At least six people have been killed in a suspected drone attack on an airport near the city of Sulaymaniyah in the semi-autonomous Kurdish region in northern Iraq, official sources have told Al Jazeera. Al Jazeera's Mahmoud Abdelwahed, reporting from the Iraqi capital Baghdad, said that the Arbat airport, located 50km (30 miles) to the east of Sulaimaniya, has been used by the "anti-terrorism" combat apparatus that is part of Sulaymaniyah security forces. "Whether all the victims are from the anti-terrorism apparatus remains to be known," he said. The airport was used for agricultural purposes in the past. Two members of the Kurdish security forces were wounded in the attack and were rushed to a military hospital in Sulaimaniya under tight security, a police source told Reuters.