Goto

Collaborating Authors

 Government


China launches lunar probe to take samples from far side of the moon

FOX News

Former National Security Adviser Robert O'Brien joins'Life, Liberty & Levin' to discuss the Biden administration's foreign policy in the Middle East. China on Friday launched a lunar probe to land on the far side of the moon and return with samples that could provide insights into differences between the less-explored region and the better-known near side. It is the latest advance in China's increasingly sophisticated space exploration program, which is now competing with the U.S., still the leader in space. China also has a three-member crew on its own orbiting space station and aims to put astronauts on the moon by 2030. Three Chinese lunar probe missions are planned over the next four years.


2 military personnel to face court martial over drone attack that killed 85 villagers in Nigeria

FOX News

Fox News State Department and foreign policy correspondent Gillian Turner has the latest on the Israel-Hamas war on'Special Report.' Two Nigerian military personnel will face a court martial over the killing of 85 villagers in a military drone attack in December in the West African nation's conflict-battered north, authorities said, prompting calls from a rights group Friday for more transparency and justice for victims. The two personnel will be subjected to military justice proceedings "for acts of omission or commission" after investigations found that the civilians killed by the strike "were mistaken for terrorists," Nigeria's Defense Headquarters spokesperson Maj. Gen. Edward Buba said in a statement Thursday without providing further details. Nigeria's military often conducts air raids as it fights the extremist violence and rebel attacks that have destabilized Nigeria's northern region for more than a decade, often leaving civilian casualties in its wake.


I Read Everything Elon Musk Posted For a Week. Send Help.

Mother Jones

Last January, not long after agreeing with an actual Nazi that western Jews have brought antisemitism upon themselves by welcoming "hordes of minorities" to their countries, Elon Musk took a quick trip to Poland. The billionaire chief of SpaceX, Tesla, and X laid a wreath at Auschwitz and then preceded on to a symposium in Krakow, where he told the conservative commentator Ben Shapiro that social media could have averted the Holocaust and bragged that he considered himself "aspirationally Jewish." The tweet, he explained in a different interview, at a different symposium "might be literally the worst and dumbest post I've ever done." But he did not take it down, nor has he moderated his views. If anything his descent into the online fever swamp has only accelerated.


Would You Still Use Google if It Didn't Pay Apple 20 Billion to Get on Your iPhone?

WIRED

Microsoft has poured over 100 billion into developing its Bing search engine over the past two decades but has little market share to show for it. About nine out of every 10 web searches in the US are made through Google, with Bing splitting the remaining queries with a long list of small competitors. On Thursday the US government asked a federal judge in Washington, DC, to rule that Google maintains that lead illegally, by unfairly manipulating users to keep Microsoft and other competitors down. Google's dominance drove the US Department of Justice to sue the company in 2020 alleging that it had violated antitrust law by using exclusionary contracts to maintain a monopoly. The two sides went into a secretive trial at the end of last year before breaking for nearly five months for US Judge Amit Mehta to digest the evidence.


Learning label-label correlations in Extreme Multi-label Classification via Label Features

arXiv.org Artificial Intelligence

Extreme Multi-label Text Classification (XMC) involves learning a classifier that can assign an input with a subset of most relevant labels from millions of label choices. Recent works in this domain have increasingly focused on a symmetric problem setting where both input instances and label features are short-text in nature. Short-text XMC with label features has found numerous applications in areas such as query-to-ad-phrase matching in search ads, title-based product recommendation, prediction of related searches. In this paper, we propose Gandalf, a novel approach which makes use of a label co-occurrence graph to leverage label features as additional data points to supplement the training distribution. By exploiting the characteristics of the short-text XMC problem, it leverages the label features to construct valid training instances, and uses the label graph for generating the corresponding soft-label targets, hence effectively capturing the label-label correlations. Surprisingly, models trained on these new training instances, although being less than half of the original dataset, can outperform models trained on the original dataset, particularly on the PSP@k metric for tail labels. With this insight, we aim to train existing XMC algorithms on both, the original and new training instances, leading to an average 5% relative improvements for 6 state-of-the-art algorithms across 4 benchmark datasets consisting of up to 1.3M labels. Gandalf can be applied in a plug-and-play manner to various methods and thus forwards the state-of-the-art in the domain, without incurring any additional computational overheads.


A Survey on Contribution Evaluation in Vertical Federated Learning

arXiv.org Artificial Intelligence

Vertical Federated Learning (VFL) has emerged as a critical approach in machine learning to address privacy concerns associated with centralized data storage and processing. VFL facilitates collaboration among multiple entities with distinct feature sets on the same user population, enabling the joint training of predictive models without direct data sharing. A key aspect of VFL is the fair and accurate evaluation of each entity's contribution to the learning process. This is crucial for maintaining trust among participating entities, ensuring equitable resource sharing, and fostering a sustainable collaboration framework. This paper provides a thorough review of contribution evaluation in VFL. We categorize the vast array of contribution evaluation techniques along the VFL lifecycle, granularity of evaluation, privacy considerations, and core computational methods. We also explore various tasks in VFL that involving contribution evaluation and analyze their required evaluation properties and relation to the VFL lifecycle phases. Finally, we present a vision for the future challenges of contribution evaluation in VFL. By providing a structured analysis of the current landscape and potential advancements, this paper aims to guide researchers and practitioners in the design and implementation of more effective, efficient, and privacy-centric VFL solutions. Relevant literature and open-source resources have been compiled and are being continuously updated at the GitHub repository: \url{https://github.com/cuiyuebing/VFL_CE}.


Fake Artificial Intelligence Generated Contents (FAIGC): A Survey of Theories, Detection Methods, and Opportunities

arXiv.org Artificial Intelligence

In recent years, generative artificial intelligence models, represented by Large Language Models (LLMs) and Diffusion Models (DMs), have revolutionized content production methods. These artificial intelligence-generated content (AIGC) have become deeply embedded in various aspects of daily life and work. However, these technologies have also led to the emergence of Fake Artificial Intelligence Generated Content (FAIGC), posing new challenges in distinguishing genuine information. It is crucial to recognize that AIGC technology is akin to a double-edged sword; its potent generative capabilities, while beneficial, also pose risks for the creation and dissemination of FAIGC. In this survey, We propose a new taxonomy that provides a more comprehensive breakdown of the space of FAIGC methods today. Next, we explore the modalities and generative technologies of FAIGC. We introduce FAIGC detection methods and summarize the related benchmark from various perspectives. Finally, we discuss outstanding challenges and promising areas for future research.


MedReadMe: A Systematic Study for Fine-grained Sentence Readability in Medical Domain

arXiv.org Artificial Intelligence

Medical texts are notoriously challenging to read. Properly measuring their readability is the first step towards making them more accessible. In this paper, we present a systematic study on fine-grained readability measurements in the medical domain at both sentence-level and span-level. We introduce a new dataset MedReadMe, which consists of manually annotated readability ratings and fine-grained complex span annotation for 4,520 sentences, featuring two novel "Google-Easy" and "Google-Hard" categories. It supports our quantitative analysis, which covers 650 linguistic features and automatic complex word and jargon identification. Enabled by our high-quality annotation, we benchmark and improve several state-of-the-art sentence-level readability metrics for the medical domain specifically, which include unsupervised, supervised, and prompting-based methods using recently developed large language models (LLMs). Informed by our fine-grained complex span annotation, we find that adding a single feature, capturing the number of jargon spans, into existing readability formulas can significantly improve their correlation with human judgments. We will publicly release the dataset and code.


Semantic Scaling: Bayesian Ideal Point Estimates with Large Language Models

arXiv.org Artificial Intelligence

This paper introduces "Semantic Scaling," a novel method for ideal point estimation from text. I leverage large language models to classify documents based on their expressed stances and extract survey-like data. I then use item response theory to scale subjects from these data. Semantic Scaling significantly improves on existing text-based scaling methods, and allows researchers to explicitly define the ideological dimensions they measure. This represents the first scaling approach that allows such flexibility outside of survey instruments and opens new avenues of inquiry for populations difficult to survey. Additionally, it works with documents of varying length, and produces valid estimates of both mass and elite ideology. I demonstrate that the method can differentiate between policy preferences and in-group/out-group affect. Among the public, Semantic Scaling out-preforms Tweetscores according to human judgement; in Congress, it recaptures the first dimension DW-NOMINATE while allowing for greater flexibility in resolving construct validity challenges.


Public-private funding models in open source software development: A case study on scikit-learn

arXiv.org Artificial Intelligence

Governments are increasingly funding open source software (OSS) development to support software security, digital sovereignty, and national competitiveness in science and innovation, amongst others. However, little is known about how OSS developers evaluate the relative benefits and drawbacks of governmental funding for OSS. This study explores this question through a case study on scikit-learn, a Python library for machine learning, funded by public research grants, commercial sponsorship, micro-donations, and a 32 euro million grant announced in France's artificial intelligence strategy. Through 25 interviews with scikit-learn's maintainers and funders, this study makes two key contributions. First, it contributes empirical findings about the benefits and drawbacks of public and private funding in an impactful OSS project, and the governance protocols employed by the maintainers to balance the diverse interests of their community and funders. Second, it offers practical lessons on funding for OSS developers, governments, and companies based on the experience of scikit-learn. The paper concludes with key recommendations for practitioners and future research directions.