Goto

Collaborating Authors

 Government


OrderSum: Semantic Sentence Ordering for Extractive Summarization

arXiv.org Artificial Intelligence

The sentence-level framework defines extractive summarization as an individual sentence selection problem, determining whether each sentence in a document should be included in the summary. However, the sentence-level framework often produces summaries that contain only general sentences or repeat important but similar sentences (Narayan et al., 2018b; Zhong et al., 2020). The summary-level framework overcomes this limitation by defining extractive summarization as a summary ranking problem rather than a sentence selection problem. The main idea of the summary-level framework is to generate a set of candidate summaries consisting of different sentences, and then rank them to select the best summary. By considering sentence composition at the entire summary level rather than sentence by sentence, this approach enables each sentence in the summary to convey different, specific information (Narayan et al., 2018b; Zhong et al., 2020). Previous work in both frameworks has primarily focused on improving which sentences to include in the summary, or in other words, sentence inclusion. However, to the best of our knowledge, the importance of sentence order in summaries has not been highlighted since the era of graph-based extractive summarization (Mihalcea and Ta-rau, 2004; Erkan and Radev, 2004). The sentence order of a text plays a crucial role not only in readability but also in its meaning (Yin et al., 2019; Lo-geswaran et al., 2018). Table 1 illustrates how the arXiv:2502.16180v1


The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination

arXiv.org Artificial Intelligence

Hallucination is a persistent challenge in large language models (LLMs), where even with rigorous quality control, models often generate distorted facts. This paradox, in which error generation continues despite high-quality training data, calls for a deeper understanding of the underlying LLM mechanisms. To address it, we propose a novel concept: knowledge overshadowing, where model's dominant knowledge can obscure less prominent knowledge during text generation, causing the model to fabricate inaccurate details. Building on this idea, we introduce a novel framework to quantify factual hallucinations by modeling knowledge overshadowing. Central to our approach is the log-linear law, which predicts that the rate of factual hallucination increases linearly with the logarithmic scale of (1) Knowledge Popularity, (2) Knowledge Length, and (3) Model Size. The law provides a means to preemptively quantify hallucinations, offering foresight into their occurrence even before model training or inference. Built on overshadowing effect, we propose a new decoding strategy CoDa, to mitigate hallucinations, which notably enhance model factuality on Overshadow (27.9%), MemoTrap (13.1%) and NQ-Swap (18.3%). Our findings not only deepen understandings of the underlying mechanisms behind hallucinations but also provide actionable insights for developing more predictable and controllable language models.


Subspace Recovery in Winsorized PCA: Insights into Accuracy and Robustness

arXiv.org Machine Learning

In this paper, we explore the theoretical properties of subspace recovery using Winsorized Principal Component Analysis (WPCA), utilizing a common data transformation technique that caps extreme values to mitigate the impact of outliers. Despite the widespread use of winsorization in various tasks of multivariate analysis, its theoretical properties, particularly for subspace recovery, have received limited attention. We provide a detailed analysis of the accuracy of WPCA, showing that increasing the number of samples while decreasing the proportion of outliers guarantees the consistency of the sample subspaces from WPCA with respect to the true population subspace. Furthermore, we establish perturbation bounds that ensure the WPCA subspace obtained from contaminated data remains close to the subspace recovered from pure data. Additionally, we extend the classical notion of breakdown points to subspace-valued statistics and derive lower bounds for the breakdown points of WPCA. Our analysis demonstrates that WPCA exhibits strong robustness to outliers while maintaining consistency under mild assumptions. A toy example is provided to numerically illustrate the behavior of the upper bounds for perturbation bounds and breakdown points, emphasizing winsorization's utility in subspace recovery.


OpenAI bans Chinese accounts using ChatGPT to edit code for social media surveillance

Engadget

OpenAI has banned the accounts of a group of Chinese users who had attempted to use ChatGPT to debug and edit code for an AI social media surveillance tool, the company said Friday. The campaign, which OpenAI calls Peer Review, saw the group prompt ChatGPT to generate sales pitches for a program those documents suggest was designed to monitor anti-Chinese sentiment on X, Facebook, YouTube, Instagram and other platforms. The operation appears to have been particularly interested in spotting calls for protests against human rights violations in China, with the intent of sharing those insights with the country's authorities. "This network consisted of ChatGPT accounts that operated in a time pattern consistent with mainland Chinese business hours, prompted our models in Chinese, and used our tools with a volume and variety consistent with manual prompting, rather than automation," said OpenAI. "The operators used our models to proofread claims that their insights had been sent to Chinese embassies abroad, and to intelligence agents monitoring protests in countries including the United States, Germany and the United Kingdom."


Elon Musk's DOGE reportedly cuts staff at agency that regulates Elon Musk's Tesla

Engadget

Elon Musk's chainsaw has been swinging through the federal government over the last few weeks, with his Department of Government Efficiency (DOGE) chopping down budgets and excising staff at a number of agencies. Among those affected is the National Highway Traffic Safety Administration (NHTSA), which is said to be losing about 10 percent of its relatively small headcount through buyouts and firings. According to The Washington Post, between 70 and 80 people are departing the agency, which is responsible for road safety in the US. Those ex-employees are said to have worked in a number of areas, such as safety grant funding and crash test dummies. The DOGE cull also impacted three people from a very small team that was working on the safety of autonomous vehicles, such as those from Alphabet's Waymo, Amazon's Zoox and -- hey, look at that! -- Elon Musk's Tesla.


China, Iran-based threat actors have found new ways to to use American AI models for covert influence: Report

FOX News

Threat actors, some likely based in China and Iran, are formulating new ways to hijack and utilize American artificial intelligence (AI) models for malicious intent, including covert influence operations, according to a new report from OpenAI. The February report includes two disruptions involving threat actors that appear to have originated from China. According to the report, these actors have used, or at least attempted to use, models built by OpenAI and Meta. In one example, OpenAI banned a ChatGPT account that generated comments critical of Chinese dissident Cai Xia. The comments were posted on social media by accounts that claimed to be people based in India and the U.S.


DOGE Sparks Surveillance Fear Across the US Government

WIRED

This month, Andrew Bernier, a US Army Corps of Engineers researcher and a union leader, says that he has received a barrage of menacing messages from the same anonymous email account. Unfolding like short chapters in a dystopian novel, they have spoken of the genius of Elon Musk, referenced the power of the billionaire's so-called Department of Government Efficiency (DOGE), and foretold the downfall of "corrupt" union bosses. But the most eerie thing about the emails, which Bernier says began arriving after he filed an official charge accusing the Trump administration of violating his union's collective bargaining agreement, is that they included personal details about his life--some of which he believes might have come from surveillance of his work laptop. The author referenced Bernier's union activities, nickname, job, travel details, and even the green notebook he regularly uses. The most recent email implied that his computer was loaded with spyware.


Black lawmakers continue push to assist descendants of slaves in California

Los Angeles Times

The California Legislative Black Caucus on Thursday proposed a package of reparations for the descendants of African Americans who were enslaved in the United States, proposals that include preferences for public university admissions and financial assistance for first-time home buyers. The package contains 15 bills in what caucus members said will be a multiyear effort to repair the generational harms and discrimination suffered by the descendants of slaves in California. In 2020, Gov. Gavin Newsom and California lawmakers formed a "first in the nation" state task force to study and propose remedies for the legacy of slavery. During the end of the legislative session last year, reform advocates were frustrated that the legislature, which was limited by a tight state budget and a high-stakes election year, passed only 10 of the 14 bills prioritized by the Legislative Black Caucus. "We are picking up where we left off last year," said Assemblymember Lori Wilson (D-Suisun City) at a press conference Thursday morning.


Japan to promote digital transformation for water systems

The Japan Times

A Japanese government panel agreed Thursday to promote digital transformation to tackle the aging of public infrastructure, including water supply and sewage systems. This followed a high-profile road collapse incident in Yashio, Saitama Prefecture, last month, which is believed to have been caused by a broken sewage pipe. At a meeting of the digital administrative and fiscal reform panel, Prime Minister Shigeru Ishiba, who heads the group, instructed related officials to urgently work on the use of digital technologies for water and sewage systems to ensure that their operations by local governments are sustainable. He called for introducing such technologies within about three years, against the previous deadline of five years. For water and sewage systems, satellites and artificial intelligence systems will be used to collect and analyze data on temperature, geology and other factors to identify areas where water leaks may occur.


Integrating Generative AI in Cybersecurity Education: Case Study Insights on Pedagogical Strategies, Critical Thinking, and Responsible AI Use

arXiv.org Artificial Intelligence

The rapid advancement of Generative Artificial Intelligence (GenAI) has introduced new opportunities for transforming higher education, particularly in fields that require analytical reasoning and regulatory compliance, such as cybersecurity management. This study presents a structured framework for integrating GenAI tools into cybersecurity education, demonstrating their role in fostering critical thinking, real-world problem-solving, and regulatory awareness. The implementation strategy followed a two-stage approach, embedding GenAI within tutorial exercises and assessment tasks. Tutorials enabled students to generate, critique, and refine AI-assisted cybersecurity policies, while assessments required them to apply AI-generated outputs to real-world scenarios, ensuring alignment with industry standards and regulatory requirements. Findings indicate that AI-assisted learning significantly enhanced students' ability to evaluate security policies, refine risk assessments, and bridge theoretical knowledge with practical application. Student reflections and instructor observations revealed improvements in analytical engagement, yet challenges emerged regarding AI over-reliance, variability in AI literacy, and the contextual limitations of AI-generated content. Through structured intervention and research-driven refinement, students were able to recognize AI strengths as a generative tool while acknowledging its need for human oversight. This study further highlights the broader implications of AI adoption in cybersecurity education, emphasizing the necessity of balancing automation with expert judgment to cultivate industry-ready professionals. Future research should explore the long-term impact of AI-driven learning on cybersecurity competency, as well as the potential for adaptive AI-assisted assessments to further personalize and enhance educational outcomes.