Goto

Collaborating Authors

 Government


State-of-the-art AI-based Learning Approaches for Deepfake Generation and Detection, Analyzing Opportunities, Threading through Pros, Cons, and Future Prospects

arXiv.org Artificial Intelligence

The rapid advancement of deepfake technologies, specifically designed to create incredibly lifelike facial imagery and video content, has ignited a remarkable level of interest and curiosity across many fields, including forensic analysis, cybersecurity and the innovative creation of digital characters. By harnessing the latest breakthroughs in deep learning methods, such as Generative Adversarial Networks, Variational Autoencoders, Few-Shot Learning Strategies, and Transformers, the outcomes achieved in generating deepfakes have been nothing short of astounding and transformative. Also, the ongoing evolution of detection technologies is being developed to counteract the potential for misuse associated with deepfakes, effectively addressing critical concerns that range from political manipulation to the dissemination of fake news and the ever-growing issue of cyberbullying. This comprehensive review paper meticulously investigates the most recent developments in deepfake generation and detection, including around 400 publications, providing an in-depth analysis of the cutting-edge innovations shaping this rapidly evolving landscape. Starting with a thorough examination of systematic literature review methodologies, we embark on a journey that delves into the complex technical intricacies inherent in the various techniques used for deepfake generation, comprehensively addressing the challenges faced, potential solutions available, and the nuanced details surrounding manipulation formulations. Subsequently, the paper is dedicated to accurately benchmarking leading approaches against prominent datasets, offering thorough assessments of the contributions that have significantly impacted these vital domains. Ultimately, we engage in a thoughtful discussion of the existing challenges, paving the way for continuous advancements in this critical and ever-dynamic study area.


Physics-informed Gaussian Processes for Safe Envelope Expansion

arXiv.org Artificial Intelligence

Flight test analysis often requires predefined test points with arbitrarily tight tolerances, leading to extensive and resource-intensive experimental campaigns. To address this challenge, we propose a novel approach to flight test analysis using Gaussian processes (GPs) with physics-informed mean functions to estimate aerodynamic quantities from arbitrary flight test data, validated using real T-38 aircraft data collected in collaboration with the United States Air Force Test Pilot School. We demonstrate our method by estimating the pitching moment coefficient without requiring predefined or repeated flight test points, significantly reducing the need for extensive experimental campaigns. Our approach incorporates aerodynamic models as priors within the GP framework, enhancing predictive accuracy across diverse flight conditions and providing robust uncertainty quantification. Key contributions include the integration of physics-based priors in a probabilistic model, which allows for precise computation from arbitrary flight test maneuvers, and the demonstration of our method capturing relevant dynamic characteristics such as short-period mode behavior. The proposed framework offers a scalable and generalizable solution for efficient data-driven flight test analysis and is able to accurately predict the short period frequency and damping for the T-38 across several Mach and dynamic pressure profiles.


Diffusion Policies for Generative Modeling of Spacecraft Trajectories

arXiv.org Artificial Intelligence

Despite its promise and the tremendous advances in nonlinear optimization solvers in recent years, trajectory optimization has primarily been constrained to offline usage due to the limited compute capabilities of radiation hardened flight computers [3]. However, with a flurry of proposed mission concepts that call for increasingly greater on-board autonomy [4], bridging this gap in the state-of-practice is necessary to allow for scaling current trajectory design techniques for future missions. Recently, researchers have turned to machine learning and data-driven techniques as a promising method for reducing the runtimes necessary for solving challenging constrained optimization problems [5, 6]. Such approaches entail learning what is known as the problem-to-solution mapping between the problem parameters that vary between repeated instances of solving the trajectory optimization problem to the full optimization solution and these works typically use a Deep Neural Network (DNN) to model this mapping [7-9]. Given parameters of new instances of the trajectory optimization problem, this problem-to-solution mapping can be used online to yield candidate trajectories to warm start the nonlinear optimization solver and this warm start can enable significant solution speed ups. One shortcoming of these aforementioned data-driven approaches is that they have limited scope of use and the learned problem-to-solution mapping only applies for one specific trajectory optimization formulation. With a change to the mission design specifications that yields, e.g., a different optimization constraint, a new problem-to-solution mapping has to be learned offline and this necessitates generating a new dataset of solved trajectory optimization problems. To this end, our work explores the use of compositional diffusion modeling to allow for generalizable learning of the problem-to-solution mapping and equip mission designers with the ability to interleave different learned models to satisfy a rich set of trajectory design specifications. Compositional diffusion modeling enables training of a model to both sample and plan from.


Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization

arXiv.org Artificial Intelligence

In the fast-changing realm of information, the capacity to construct coherent timelines from extensive event-related content has become increasingly significant and challenging. The complexity arises in aggregating related documents to build a meaningful event graph around a central topic. This paper proposes CHRONOS - Causal Headline Retrieval for Open-domain News Timeline SummarizatiOn via Iterative Self-Questioning, which offers a fresh perspective on the integration of Large Language Models (LLMs) to tackle the task of Timeline Summarization (TLS). By iteratively reflecting on how events are linked and posing new questions regarding a specific news topic to gather information online or from an offline knowledge base, LLMs produce and refresh chronological summaries based on documents retrieved in each round. Furthermore, we curate Open-TLS, a novel dataset of timelines on recent news topics authored by professional journalists to evaluate open-domain TLS where information overload makes it impossible to find comprehensive relevant documents from the web. Our experiments indicate that CHRONOS is not only adept at open-domain timeline summarization, but it also rivals the performance of existing state-of-the-art systems designed for closed-domain applications, where a related news corpus is provided for summarization.


Agentic Systems: A Guide to Transforming Industries with Vertical AI Agents

arXiv.org Artificial Intelligence

The evolution of agentic systems represents a significant milestone in artificial intelligence and modern software systems, driven by the demand for vertical intelligence tailored to diverse industries. These systems enhance business outcomes through adaptability, learning, and interaction with dynamic environments. At the forefront of this revolution are Large Language Model (LLM) agents, which serve as the cognitive backbone of these intelligent systems. In response to the need for consistency and scalability, this work attempts to define a level of standardization for Vertical AI agent design patterns by identifying core building blocks and proposing a \textbf{Cognitive Skills } Module, which incorporates domain-specific, purpose-built inference capabilities. Building on these foundational concepts, this paper offers a comprehensive introduction to agentic systems, detailing their core components, operational patterns, and implementation strategies. It further explores practical use cases and examples across various industries, highlighting the transformative potential of LLM agents in driving industry-specific applications.


What is a Social Media Bot? A Global Comparison of Bot and Human Characteristics

arXiv.org Artificial Intelligence

Chatter on social media is 20% bots and 80% humans. Chatter by bots and humans is consistently different: bots tend to use linguistic cues that can be easily automated while humans use cues that require dialogue understanding. Bots use words that match the identities they choose to present, while humans may send messages that are not related to the identities they present. Bots and humans differ in their communication structure: sampled bots have a star interaction structure, while sampled humans have a hierarchical structure. These conclusions are based on a large-scale analysis of social media tweets across ~200mil users across 7 events. Social media bots took the world by storm when social-cybersecurity researchers realized that social media users not only consisted of humans but also of artificial agents called bots. These bots wreck havoc online by spreading disinformation and manipulating narratives. Most research on bots are based on special-purposed definitions, mostly predicated on the event studied. This article first begins by asking, "What is a bot?", and we study the underlying principles of how bots are different from humans. We develop a first-principle definition of a social media bot. With this definition as a premise, we systematically compare characteristics between bots and humans across global events, and reflect on how the software-programmed bot is an Artificial Intelligent algorithm, and its potential for evolution as technology advances. Based on our results, we provide recommendations for the use and regulation of bots. Finally, we discuss open challenges and future directions: Detect, to systematically identify these automated and potentially evolving bots; Differentiate, to evaluate the goodness of the bot in terms of their content postings and relationship interactions; Disrupt, to moderate the impact of malicious bots.


Beyond Static Datasets: A Behavior-Driven Entity-Specific Simulation to Overcome Data Scarcity and Train Effective Crypto Anti-Money Laundering Models

arXiv.org Artificial Intelligence

For different factors/reasons, ranging from inherent characteristics and features providing decentralization, enhanced privacy, ease of transactions, etc., to implied external hardships in enforcing regulations, contradictions in data sharing policies, etc., cryptocurrencies have been severely abused for carrying out numerous malicious and illicit activities including money laundering, darknet transactions, scams, terrorism financing, arm trades. However, money laundering is a key crime to be mitigated to also suspend the movement of funds from other illicit activities. Billions of dollars are annually being laundered. It is getting extremely difficult to identify money laundering in crypto transactions owing to many layering strategies available today, and rapidly evolving tactics, and patterns the launderers use to obfuscate the illicit funds. Many detection methods have been proposed ranging from naive approaches involving complete manual investigation to machine learning models. However, there are very limited datasets available for effectively training machine learning models. Also, the existing datasets are static and class-imbalanced, posing challenges for scalability and suitability to specific scenarios, due to lack of customization to varying requirements. This has been a persistent challenge in literature. In this paper, we propose behavior embedded entity-specific money laundering-like transaction simulation that helps in generating various transaction types and models the transactions embedding the behavior of several entities observed in this space. The paper discusses the design and architecture of the simulator, a custom dataset we generated using the simulator, and the performance of models trained on this synthetic data in detecting real addresses involved in money laundering.


The most important tech stories of 2024, and also my favorite ones

The Guardian

Last week, we looked back at how 2024 made Elon Musk the world's most powerful man. Today, we're looking at a few other important themes that will influence the online and offline worlds in 2025. Google: Ruled an illegal monopoly in August, Google could be broken up. The results are anybody's guess, but what seemed impossible for a company worth 2.5tn is at play. The US has asked the judge in the case for a wholesale breakup of the giant, which would force it to divest Chrome, the world's most popular browser and one of Google's core businesses.


22 health care predictions for 2025 from medical researchers

FOX News

First, the integration of artificial intelligence-facilitated algorithms for the early detection of cardiovascular illness, which will move us closer toward early prevention. We also envision a focus on using genetically informed treatments to reduce the risk of atherosclerotic heart disease, valvular heart disease and heart failure. Together, these important advances will usher in an era of personalized health care in cardiovascular disease."


Learning Curve: The new players in Congress

FOX News

Fox News senior congressional correspondent Chad Pergram joins'Fox News Live' to explain how he prepares to report on Congress for the upcoming year. Every two years, the period between the November election and when the new Congress begins is often the busiest swath of time for covering Congress. Reporters are trying to figure out who won their elections and who lost. The existing Congress is back, attempting to prevent a government shutdown and often plowing through a landscape of other major legislation. There are often leadership elections.