Government
Beneath Surface Similarity: Large Language Models Make Reasonable Scientific Analogies after Structure Abduction
Yuan, Siyu, Chen, Jiangjie, Ge, Xuyang, Xiao, Yanghua, Yang, Deqing
The vital role of analogical reasoning in human cognition allows us to grasp novel concepts by linking them with familiar ones through shared relational structures. Despite the attention previous research has given to word analogies, this work suggests that Large Language Models (LLMs) often overlook the structures that underpin these analogies, raising questions about the efficacy of word analogies as a measure of analogical reasoning skills akin to human cognition. In response to this, our paper introduces a task of analogical structure abduction, grounded in cognitive psychology, designed to abduce structures that form an analogy between two systems. In support of this task, we establish a benchmark called SCAR, containing 400 scientific analogies from 13 distinct fields, tailored for evaluating analogical reasoning with structure abduction. The empirical evidence underlines the continued challenges faced by LLMs, including ChatGPT and GPT-4, in mastering this task, signifying the need for future exploration to enhance their abilities.
Statistical properties and privacy guarantees of an original distance-based fully synthetic data generation method
Chapelle, Rémy, Falissard, Bruno
Introduction: The amount of data generated by original research is growing exponentially. Publicly releasing them is recommended to comply with the Open Science principles. However, data collected from human participants cannot be released as-is without raising privacy concerns. Fully synthetic data represent a promising answer to this challenge. This approach is explored by the French Centre de Recherche en {\'E}pid{\'e}miologie et Sant{\'e} des Populations in the form of a synthetic data generation framework based on Classification and Regression Trees and an original distance-based filtering. The goal of this work was to develop a refined version of this framework and to assess its risk-utility profile with empirical and formal tools, including novel ones developed for the purpose of this evaluation.Materials and Methods: Our synthesis framework consists of four successive steps, each of which is designed to prevent specific risks of disclosure. We assessed its performance by applying two or more of these steps to a rich epidemiological dataset. Privacy and utility metrics were computed for each of the resulting synthetic datasets, which were further assessed using machine learning approaches.Results: Computed metrics showed a satisfactory level of protection against attribute disclosure attacks for each synthetic dataset, especially when the full framework was used. Membership disclosure attacks were formally prevented without significantly altering the data. Machine learning approaches showed a low risk of success for simulated singling out and linkability attacks. Distributional and inferential similarity with the original data were high with all datasets.Discussion: This work showed the technical feasibility of generating publicly releasable synthetic data using a multi-step framework. Formal and empirical tools specifically developed for this demonstration are a valuable contribution to this field. Further research should focus on the extension and validation of these tools, in an effort to specify the intrinsic qualities of alternative data synthesis methods.Conclusion: By successfully assessing the quality of data produced using a novel multi-step synthetic data generation framework, we showed the technical and conceptual soundness of the Open-CESP initiative, which seems ripe for full-scale implementation.
Discovering Mixtures of Structural Causal Models from Time Series Data
Varambally, Sumanth, Ma, Yi-An, Yu, Rose
In fields such as finance, climate science, and neuroscience, inferring causal relationships from time series data poses a formidable challenge. While contemporary techniques can handle nonlinear relationships between variables and flexible noise distributions, they rely on the simplifying assumption that data originates from the same underlying causal model. In this work, we relax this assumption and perform causal discovery from time series data originating from mixtures of different causal models. We infer both the underlying structural causal models and the posterior probability for each sample belonging to a specific mixture component. Our approach employs an end-to-end training process that maximizes an evidence-lower bound for data likelihood. Through extensive experimentation on both synthetic and real-world datasets, we demonstrate that our method surpasses state-of-the-art benchmarks in causal discovery tasks, particularly when the data emanates from diverse underlying causal graphs. Theoretically, we prove the identifiability of such a model under some mild assumptions.
The Bayesian Context Trees State Space Model for time series modelling and forecasting
Papageorgiou, Ioannis, Kontoyiannis, Ioannis
A hierarchical Bayesian framework is introduced for developing rich mixture models for real-valued time series, partly motivated by important applications in financial time series analysis. At the top level, meaningful discrete states are identified as appropriately quantised values of some of the most recent samples. These observable states are described as a discrete context-tree model. At the bottom level, a different, arbitrary model for real-valued time series -- a base model -- is associated with each state. This defines a very general framework that can be used in conjunction with any existing model class to build flexible and interpretable mixture models. We call this the Bayesian Context Trees State Space Model, or the BCT-X framework. Efficient algorithms are introduced that allow for effective, exact Bayesian inference and learning in this setting; in particular, the maximum a posteriori probability (MAP) context-tree model can be identified. These algorithms can be updated sequentially, facilitating efficient online forecasting. The utility of the general framework is illustrated in two particular instances: When autoregressive (AR) models are used as base models, resulting in a nonlinear AR mixture model, and when conditional heteroscedastic (ARCH) models are used, resulting in a mixture model that offers a powerful and systematic way of modelling the well-known volatility asymmetries in financial data. In forecasting, the BCT-X methods are found to outperform state-of-the-art techniques on simulated and real-world data, both in terms of accuracy and computational requirements. In modelling, the BCT-X structure finds natural structure present in the data. In particular, the BCT-ARCH model reveals a novel, important feature of stock market index data, in the form of an enhanced leverage effect.
An Alleged Deepfake of UK Opposition Leader Keir Starmer Shows the Dangers of Fake Audio
As members of the UK's largest opposition party gathered in Liverpool for their party conference--probably their last before the UK holds a general election--a potentially explosive audio file started circulating on X, formerly known as Twitter. The 25-second recording was posted by an X account with the handle "@Leo_Hutz" that was set up in January 2023. In the clip, Sir Keir Starmer, the Labour Party leader, is apparently heard swearing repeatedly at a staffer. "I have obtained audio of Keir Starmer verbally abusing his staffers at [the Labour Party] conference," the X account posted. "This disgusting bully is about to become our next PM."
How to fight for internet freedom
Last week, Freedom House, a human rights advocacy group, released its annual review of the state of internet freedom around the world; it's one of the most important trackers out there if you want to understand changes to digital free expression. As I wrote, the report shows that generative AI is already a game changer in geopolitics. Globally, internet freedom has never been lower, and the number of countries that have blocked websites for political, social, and religious speech has never been higher. Also, the number of countries that arrested people for online expression reached a record high. These issues are particularly urgent before we head into a year with over 50 elections worldwide; as Freedom House has noted, election cycles are times when internet freedom is often most under threat.
Democrats indicated they'd help save McCarthy before voting to oust him: sources
Rep. Carlos Gimenez, R-Fla., joined'Fox & Friends Weekend' to discuss Israel's decision to declare war and the GOP's efforts to determine a new House Speaker. Multiple House Democrats had indicated to GOP lawmakers that they would help former speaker Kevin McCarthy, R-Calif., avoid being ousted on Tuesday, two sources told Fox News Digital. A GOP member of the Problem Solvers Caucus indicated that right up until the final days, Democrats signaled they may at least be open to voting "present" to lower the threshold needed for McCarthy's political survival. The lawmaker pointed to Rep. Matt Cartwright, D-Pa., who is not a member of the Problem Solvers Caucus, but suggested early on that he could be open to helping McCarthy. "Even people like that were saying they were going to vote present. And something changed over the weekend. So yes, the members of the Problem Solvers gave absolutely no indication that they were going to side with [Rep. Gaetz introduced the motion to vacate against McCarthy on Monday evening. The next day, seven other Republicans joined him and every House Democrat to oust McCarthy from leadership. Acknowledging that Gaetz would likely pull the move again if it failed the first time, the Republican who spoke with Fox News Digital said he and other GOP Problem Solvers appealed to Democrats to vote "present" on the initial procedural vote in order to buy time to pull together a bipartisan proposal on a House Rules overhaul, which would have likely made it harder for members to topple the speaker. "We wanted them to vote present for the first round on the motion, to make the motion to table, so that they could have time to rewrite the rules package.
The US, not China, should take the lead on AI
Senior fellow at the Gatestone Institute Gordon Chang joined'Cavuto Live' to discuss the U.S.'s relationship with China amid the highly anticipated G20 Summit. Emerging technologies like artificial intelligence (AI) should be used as "tools of opportunity, not as weapons of oppression," President Biden remarked recently. But this exhortation makes his subsequent vow to work directly with "our competitors" to harness the power of AI "for good" all the more curious. Working with our competitors, like China, would only empower the Chinese Communist Party (CCP) to write the rules of the road for AI. And we don't want China in the driver's seat.
G7 to draw up AI code of conduct this autumn: Kishida
Prime Minister Fumio Kishida unveiled a plan on Monday to hold a video conference with Group of Seven leaders this autumn to formulate international guidelines and a code of conduct for developers of artificial intelligence (AI) tools. Kishida showed the plan in a speech at a special session of the U.N.-sponsored Internet Governance Forum in Kyoto. The theme of the guidelines and code of conduct is part of the Hiroshima AI Process, an initiative for international best practices regarding generative AI, according to the Japanese leader. Kishida also said that the Japanese government's new economic package, planned to be drawn up late this month, will include aid for the development of computational resources, used for processing huge volumes of data needed for AI development and use, and of basic computational models, as well as stepping up the introduction of AI in small businesses and the medical field. The Hiroshima AI Process, which was agreed on at the G7 summit held in Hiroshima in May, also calls for creating international guidelines by the end of the year that will also cover generative AI users.
Russia-Ukraine war: List of key events, day 593
Air Force spokesperson Yuriy Ihnat said that Ukraine was expecting a record number of Russian drone attacks this winter. Ihnat told national television that data Russia had already used a "record' number of more than 500 Iranian-made Shahed drones in September, compared with about 1,000 over a six-month period during last winter. Governor Oleksandr Prokudin said the southern Ukrainian region of Kherson had "another terrible night" as it was targeted in some 59 Russian attacks that left 12 people injured, including a mother and her nine-month-old baby. Several houses and gas pipelines were also damaged. Four people including a nine-year-old girl were injured in a rocket attack on Konstantinivka, according to the Donetsk regional Governor Ihor Moroz. Several homes and other buildings were also damaged. The General Staff of Ukraine's Armed Forces said the situation on the battlefield in the east and south of the country remained difficult, with troops coming under intense artillery and mortar fire in and around the front line in areas including Bakhmut, Kupiansk and Lyman. The General Staff said Ukrainian forces had inflicted casualties and equipment losses on the Russians. Air Force spokesperson Yuriy Ihnat said that Ukraine was expecting a record number of Russian drone attacks this winter. Ihnat told national television that data Russia had already used a "record' number of more than 500 Iranian-made Shahed drones in September, compared with about 1,000 over a six-month period during last winter.