Goto

Collaborating Authors

 Government


Revisit Non-parametric Two-sample Testing as a Semi-supervised Learning Problem

arXiv.org Machine Learning

Learning effective data representations is crucial in answering if two samples X and Y are from the same distribution (a.k.a. the non-parametric two-sample testing problem), which can be categorized into: i) learning discriminative representations (DRs) that distinguish between two samples in a supervised-learning paradigm, and ii) learning inherent representations (IRs) focusing on data's inherent features in an unsupervised-learning paradigm. However, both paradigms have issues: learning DRs reduces the data points available for the two-sample testing phase, and learning purely IRs misses discriminative cues. To mitigate both issues, we propose a novel perspective to consider non-parametric two-sample testing as a semi-supervised learning (SSL) problem, introducing the SSL-based Classifier Two-Sample Test (SSL-C2ST) framework. While a straightforward implementation of SSL-C2ST might directly use existing state-of-the-art (SOTA) SSL methods to train a classifier with labeled data (with sample indexes X or Y) and unlabeled data (the remaining ones in the two samples), conventional two-sample testing data often exhibits substantial overlap between samples and violates SSL methods' assumptions, resulting in low test power. Therefore, we propose a two-step approach: first, learn IRs using all data, then fine-tune IRs with only labelled data to learn DRs, which can both utilize information from whole dataset and adapt the discriminative power to the given data. Extensive experiments and theoretical analysis demonstrate that SSL-C2ST outperforms traditional C2ST by effectively leveraging unlabeled data. We also offer a stronger empirically designed test achieving the SOTA performance in many two-sample testing datasets.


Japan's wealthy fail to declare 65.5 billion in income

The Japan Times

Wealthy people in Japan failed to declare a total of 65.5 billion in taxable income in the year through June, down 33.2% from the year before, a National Tax Agency report showed Friday. During the year, the agency conducted 2,407 investigations targeting the wealthy, including those with significant holdings in securities and real estate, down 18.2%. It collected back taxes totaling 17 billion, down 7.1%. Undeclared income among all people subject to investigations, including the wealthy, rose 10.2% to a record 996.4 billion. Total back taxes grew 2.2% to 139.8 billion, also a record high.


Tesla owners turn against Musk: 'I'm embarrassed driving this car around'

The Guardian

As Elon Musk has embraced Donald Trump and various far-right conspiracy theories, he has left behind an aghast cohort of Tesla owners who suddenly feel embarrassed by their own cars. Many of them are now publicly displaying their dismay at Musk on their vehicles. Sales of anti-Musk stickers have boomed since the world's richest man declared his support for Trump and helped propel him to victory in the US presidential election, as owners of Teslas, the car brand headed by Musk, try to distance themselves from the South African-born multibillionaire. The day after the election was the biggest day ever," said Matt Hiller, a Hawaii-based aquarium worker who sells a range of stickers online that denounce Musk. "People saw a billionaire supervillain buy his way into the administration and it rubbed them the wrong way." Hiller started the sticker range last year after deciding against buying a Tesla due to Musk's "amplifying of horrible people and silencing of others" on X, formerly Twitter, another of his companies. Several hundred stickers a day are now being sold, primarily to Tesla owners, Hiller said, bearing texts such as "Anti Elon Tesla Club" or "I Bought This Before Elon Went Crazy", or a picture of Musk in clown makeup with the words "Space Clown". "People keep telling me that they feel they can drive their Teslas again with these stickers," said Hiller, who has had to set aside part of his house to accommodate the growing operation. Hiller devises slogans such as "Elon Ate My Cat", a reference to a debunked falsehood about migrants eating pets in Ohio, that are then sold on Etsy and Amazon. It's a relief really to see they are awake," he said of the surging demand.


The US Army's Vision of Soldiers in Exoskeletons Lives On

WIRED

After decades of research and development, the United States Army is taking yet another run at developing a powered exoskeleton to help soldiers carry heavy loads on the battlefield--but don't expect a futuristic suit of combat armor straight out of Starship Troopers or Iron Man anytime soon. Soldiers assigned to the Army's 1-78 Field Artillery Battalion training unit at Fort Sill, Oklahoma, recently completed a three-day "proof of concept" evaluation of several off-the-shelf "exoskeleton suits" in late September and early October, officials confirmed to WIRED. The evaluation was overseen by the service's Combat Capabilities Development Command (DEVCOM), the organization responsible for developing new technology for soldiers. Official photos from the evaluation published to social media showed Advanced Individual Training students hauling artillery shells to and from a M109 Paladin self-propelled howitzer and M777-towed howitzer with telltale black exoskeleton harnesses contrasted against their camouflage uniforms, part of a field exercise undertaken "to assess the potential of human augmentation, improve soldier performance, and determine if these exoskeletons meet the demands of our warfighters," as the service put it. While a DEVCOM spokesperson declined to identify which commercially produced systems were evaluated by soldiers, the Army announced its intent in August to award a contract to exoskeleton maker SUITX to "give users experience of advanced soldier augmentation technologies," according to a government notice.


Digital Twin in Industries: A Comprehensive Survey

arXiv.org Artificial Intelligence

Industrial networks are undergoing rapid transformation driven by the convergence of emerging technologies that are revolutionizing conventional workflows, enhancing operational efficiency, and fundamentally redefining the industrial landscape across diverse sectors. Amidst this revolution, Digital Twin (DT) emerges as a transformative innovation that seamlessly integrates real-world systems with their virtual counterparts, bridging the physical and digital realms. In this article, we present a comprehensive survey of the emerging DT-enabled services and applications across industries, beginning with an overview of DT fundamentals and its components to a discussion of key enabling technologies for DT. Different from literature works, we investigate and analyze the capabilities of DT across a wide range of industrial services, including data sharing, data offloading, integrated sensing and communication, content caching, resource allocation, wireless networking, and metaverse. In particular, we present an in-depth technical discussion of the roles of DT in industrial applications across various domains, including manufacturing, healthcare, transportation, energy, agriculture, space, oil and gas, as well as robotics. Throughout the technical analysis, we delve into real-time data communications between physical and virtual platforms to enable industrial DT networking. Subsequently, we extensively explore and analyze a wide range of major privacy and security issues in DT-based industry. Taxonomy tables and the key research findings from the survey are also given, emphasizing important insights into the significance of DT in industries. Finally, we point out future research directions to spur further research in this promising area.


Generative AI Literacy: Twelve Defining Competencies

arXiv.org Artificial Intelligence

This paper introduces a competency-based model for generative artificial intelligence (AI) literacy covering essential skills and knowledge areas necessary to interact with generative AI. The competencies range from foundational AI literacy to prompt engineering and programming skills, including ethical and legal considerations. These twelve competencies offer a framework for individuals, policymakers, government officials, and educators looking to navigate and take advantage of the potential of generative AI responsibly. Embedding these competencies into educational programs and professional training initiatives can equip individuals to become responsible and informed users and creators of generative AI. The competencies follow a logical progression and serve as a roadmap for individuals seeking to get familiar with generative AI and for researchers and policymakers to develop assessments, educational programs, guidelines, and regulations.


INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge

arXiv.org Artificial Intelligence

The performance differential of large language models (LLM) between languages hinders their effective deployment in many regions, inhibiting the potential economic and societal value of generative AI tools in many communities. However, the development of functional LLMs in many languages (i.e., multilingual LLMs) is bottlenecked by the lack of high-quality evaluation resources in languages other than English. Moreover, current practices in multilingual benchmark construction often translate English resources, ignoring the regional and cultural knowledge of the environments in which multilingual systems would be used. In this work, we construct an evaluation suite of 197,243 QA pairs from local exam sources to measure the capabilities of multilingual LLMs in a variety of regional contexts. The rapid advancement of AI technologies underscores the importance of developing LLMs that are proficient across diverse linguistic and cultural contexts, ensuring fair and equitable performance for stakeholders from various language groups. However, the lack of high-quality evaluation benchmarks in many languages discourages practitioners from training multilingual LLMs to meet this challenge. This evaluation gap limits the effective deployment of LLMs for many regions, exacerbates digital divides, and inhibits the economic and societal value of AI tools in many underserved communities. The source of this gap is the multitude of challenges in evaluating LLMs for multilingual contexts. First, at a meta-level, the majority of benchmarks for LLMs are only in English (Hendrycks et al., 2020, inter alia). Technical challenges also abound due to the manner in which multilingual datasets are often collected. Certain datasets are constructed using manually applied templates, resulting in low prompt and completion diversity (Muennighoff et al., 2022). Many more are composed of translations from high-resource languages (e.g., English; Holtermann et al., 2024; Myung et al., 2024; Lai et al., 2023; Foroutan et al., 2023). These datasets often contain errors (Ponti et al., 2020; Plaza et al., 2024) and create translationese artifacts (Vanmassenhove et al., 2021; Hartung et al., 2023; Savoldi et al., 2021; Ji et al., 2023).


Handling irresolvable conflicts in the Semantic Web: an RDF-based conflict-tolerant version of the Deontic Traditional Scheme

arXiv.org Artificial Intelligence

This paper presents a new ontology that implements the well-known Deontic Traditional Scheme in RDFs and SPARQL, fit to handle irresolvable conflicts, i.e., situations in which two or more statements prescribe conflicting obligations, prohibitions, or permissions, with none of them being "stronger" than the other one(s). In our view, this paper marks a significant advancement in standard theoretical research in formal Deontic Logic. Most contemporary approaches in this field are confined to the propositional level, mainly focus on the notion of obligation, and lack implementations. The proposed framework is encoded in RDF, which is not only a first-order language but also the most widely used knowledge representation language, as it forms the foundation of the Semantic Web. Moreover, the proposed computational ontology formalizes all deontic modalities defined in the Deontic Traditional Scheme, without specifically focusing on obligations, and offers constructs to model and reason with various types of irresolvable conflicts, violations, and the interaction between deontic modalities and contextual constraints in a given state of affairs. To the best of our knowledge, no existing approach in the literature addresses all these aspects within a unified integrated framework. All examples presented and discussed in this paper, together with Java code and clear instructions to re-execute them locally, are available at https://github.com/liviorobaldo/conflict-tolerantDeonticTraditionalScheme


An AI-Driven Data Mesh Architecture Enhancing Decision-Making in Infrastructure Construction and Public Procurement

arXiv.org Artificial Intelligence

Infrastructure construction, often dubbed an "industry of industries," is closely linked with government spending and public procurement, offering significant opportunities for improved efficiency and productivity through better transparency and information access. By leveraging these opportunities, we can achieve notable gains in productivity, cost savings, and broader economic benefits. Our approach introduces an integrated software ecosystem utilizing Data Mesh and Service Mesh architectures. This system includes the largest training dataset for infrastructure and procurement, encompassing over 100 billion tokens, scientific publications, activities, and risk data, all structured by a systematic AI framework. Supported by a Knowledge Graph linked to domain-specific multi-agent tasks and Q&A capabilities, our platform standardizes and ingests diverse data sources, transforming them into structured knowledge. Leveraging large language models (LLMs) and automation, our system revolutionizes data structuring and knowledge creation, aiding decision-making in early-stage project planning, detailed research, market trend analysis, and qualitative assessments. Its web-scalable architecture delivers domain-curated information, enabling AI agents to facilitate reasoning and manage uncertainties, while preparing for future expansions with specialized agents targeting particular challenges. This integration of AI with domain expertise not only boosts efficiency and decision-making in construction and infrastructure but also establishes a framework for enhancing government efficiency and accelerating the transition of traditional industries to digital workflows. This work is poised to significantly influence AI-driven initiatives in this sector and guide best practices in AI Operations.


Noncommutative Model Selection and the Data-Driven Estimation of Real Cohomology Groups

arXiv.org Artificial Intelligence

We propose three completely data-driven methods for estimating the real cohomology groups $H^k (X ; \mathbb{R})$ of a compact metric-measure space $(X, d_X, \mu_X)$ embedded in a metric-measure space $(Y,d_Y,\mu_Y)$, given a finite set of points $S$ sampled from a uniform distrbution $\mu_X$ on $X$, possibly corrupted with noise from $Y$. We present the results of several computational experiments in the case that $X$ is embedded in $\mathbb{R}^n$, where two of the three algorithms performed well.