Government
Three armed drones intercepted and shot down near US base in northern Iraq
Senior foreign affairs correspondent Greg Palkot provides details on the major strike on an Iraqi militia leader and the U.S.'s response to Houthi attacks in the Red Sea Three armed drones were shot down in Iraq on Tuesday, near where U.S. and other international forces are stationed, officials said. Iraqi Kurdistan's counter-terrorism service said its forces intercepted and shot down the drones over Erbil airport in northern Iraq at around 5:05 a.m. It did not say if there were any casualties or damage to infrastructure. There was no immediate claim of responsibility. Similar previous attacks have been claimed by a group called the Islamic Resistance in Iraq, an umbrella group of Iran-aligned Iraqi militias.
French surveillance flights keep close watch on Russia and Ukraine, drawing boundary in European skies
A huge fire tore through an online retailer's warehouse in St. Petersburg Saturday with video showing intense flames and thick black smoke rising into the sky (CREDIT: Reuters). Seen from up here, in the cockpit of a French air force surveillance plane flying over neighboring Romania, the snow-dusted landscapes look deceptively peaceful. The dead from Russia's war, the shattered Ukrainian towns and mangled battlefields, aren't visible to the naked eye through the clouds. But French military technicians riding farther back in the aircraft, monitoring screens that display the word "secret" when idle, have a far more penetrating view. With a powerful radar that rotates six times every minute on the fuselage and a bellyful of surveillance gear, the plane can spot missile launches, airborne bombing runs and other military activity in the conflict.
What doom loop? With AI, a 'spirit of optimism' returns to San Francisco start-ups
Far from the palm trees of Miami or Austin's taco trucks, Catalin Voss has headquartered his literacy start-up between a cannabis club and pawn shop in the heart of the Mission District. Voss rents a nondescript office building in one of San Francisco's most vibrant neighborhoods as a home base for Ello, a company he co-founded in 2020 that uses speech recognition technology, powered by artificial intelligence, to help struggling students develop their reading skills. The office is within walking distance of his Noe Valley apartment and only steps away from some of the city's best taquerias and cocktail bars. And those are just a few of the perks he recited in explaining why he is headquartered in San Francisco. Voss is part of a sizable cohort of San Francisco loyalists -- old and new -- who say they are flummoxed by the "all is lost" narrative propagated by conservative media hosts and more recently a vocal contingent of tech leaders that includes billionaire entrepreneur-turned-agitator Elon Musk.
OpenAI lays out its misinformation strategy ahead of 2024 elections
As the US gears up for the 2024 presidential election, OpenAI shares its plans on suppressing misinformation related to elections worldwide, with a focus set on boosting the transparency around the origin of information. One such highlight is the use of cryptography -- as standardized by the Coalition for Content Provenance and Authenticity -- to encode the provenance of images generated by DALL-E 3. This will allow the platform to better detect AI-generated images using a provenance classifier, in order to help voters assess the reliability of certain content. This approach is similar to, if not better than, DeepMind's SynthID for digitally watermark AI-generated images and audio, as part of Google's own election content strategy published last month. Meta's AI image generator also adds an invisible watermark to its content, though the company has yet to share its readiness on tackling election-related misinformation.
Glitter or Gold? Deriving Structured Insights from Sustainability Reports via Large Language Models
Bronzini, Marco, Nicolini, Carlo, Lepri, Bruno, Passerini, Andrea, Staiano, Jacopo
Over the last decade, several regulatory bodies have started requiring the disclosure of non-financial information from publicly listed companies, in light of the investors' increasing attention to Environmental, Social, and Governance (ESG) issues. Publicly released information on sustainability practices is often disclosed in diverse, unstructured, and multi-modal documentation. This poses a challenge in efficiently gathering and aligning the data into a unified framework to derive insights related to Corporate Social Responsibility (CSR). Thus, using Information Extraction (IE) methods becomes an intuitive choice for delivering insightful and actionable data to stakeholders. In this study, we employ Large Language Models (LLMs), In-Context Learning, and the Retrieval-Augmented Generation (RAG) paradigm to extract structured insights related to ESG aspects from companies' sustainability reports. We then leverage graph-based representations to conduct statistical analyses concerning the extracted insights. These analyses revealed that ESG criteria cover a wide range of topics, exceeding 500, often beyond those considered in existing categorizations, and are addressed by companies through a variety of initiatives. Moreover, disclosure similarities emerged among companies from the same region or sector, validating ongoing hypotheses in the ESG literature. Lastly, by incorporating additional company attributes into our analyses, we investigated which factors impact the most on companies' ESG ratings, showing that ESG disclosure affects the obtained ratings more than other financial or company data.
A Framework for Scalable Ambient Air Pollution Concentration Estimation
Berrisford, Liam J, Neal, Lucy S, Buttery, Helen J, Evans, Benjamin R, Menezes, Ronaldo
Ambient air pollution remains a critical issue in the United Kingdom, where data on air pollution concentrations form the foundation for interventions aimed at improving air quality. However, the current air pollution monitoring station network in the UK is characterized by spatial sparsity, heterogeneous placement, and frequent temporal data gaps, often due to issues such as power outages. We introduce a scalable data-driven supervised machine learning model framework designed to address temporal and spatial data gaps by filling missing measurements. This approach provides a comprehensive dataset for England throughout 2018 at a 1kmx1km hourly resolution. Leveraging machine learning techniques and real-world data from the sparsely distributed monitoring stations, we generate 355,827 synthetic monitoring stations across the study area, yielding data valued at approximately \pounds70 billion. Validation was conducted to assess the model's performance in forecasting, estimating missing locations, and capturing peak concentrations. The resulting dataset is of particular interest to a diverse range of stakeholders engaged in downstream assessments supported by outdoor air pollution concentration data for NO2, O3, PM10, PM2.5, and SO2. This resource empowers stakeholders to conduct studies at a higher resolution than was previously possible.
The Impact of Differential Feature Under-reporting on Algorithmic Fairness
Akpinar, Nil-Jana, Lipton, Zachary C., Chouldechova, Alexandra
Predictive risk models in the public sector are commonly developed using administrative data that is more complete for subpopulations that more greatly rely on public services. In the United States, for instance, information on health care utilization is routinely available to government agencies for individuals supported by Medicaid and Medicare, but not for the privately insured. Critiques of public sector algorithms have identified such differential feature under-reporting as a driver of disparities in algorithmic decision-making. Yet this form of data bias remains understudied from a technical viewpoint. While prior work has examined the fairness impacts of additive feature noise and features that are clearly marked as missing, the setting of data missingness absent indicators (i.e. differential feature under-reporting) has been lacking in research attention. In this work, we present an analytically tractable model of differential feature under-reporting which we then use to characterize the impact of this kind of data bias on algorithmic fairness. We demonstrate how standard missing data methods typically fail to mitigate bias in this setting, and propose a new set of methods specifically tailored to differential feature under-reporting. Our results show that, in real world data settings, under-reporting typically leads to increasing disparities. The proposed solution methods show success in mitigating increases in unfairness.
Robust Anomaly Detection for Particle Physics Using Multi-Background Representation Learning
Gandrakota, Abhijith, Zhang, Lily, Puli, Aahlad, Cranmer, Kyle, Ngadiuba, Jennifer, Ranganath, Rajesh, Tran, Nhan
Anomaly, or out-of-distribution, detection is a promising tool for aiding discoveries of new particles or processes in particle physics. In this work, we identify and address two overlooked opportunities to improve anomaly detection for high-energy physics. First, rather than train a generative model on the single most dominant background process, we build detection algorithms using representation learning from multiple background types, thus taking advantage of more information to improve estimation of what is relevant for detection. Second, we generalize decorrelation to the multi-background setting, thus directly enforcing a more complete definition of robustness for anomaly detection. We demonstrate the benefit of the proposed robust multi-background anomaly detection algorithms on a high-dimensional dataset of particle decays at the Large Hadron Collider.
Bag of Tricks to Boost Adversarial Transferability
Zhang, Zeliang, Zhu, Rongyi, Yao, Wei, Wang, Xiaosen, Xu, Chenliang
Deep neural networks are widely known to be vulnerable to adversarial examples. However, vanilla adversarial examples generated under the white-box setting often exhibit low transferability across different models. Since adversarial transferability poses more severe threats to practical applications, various approaches have been proposed for better transferability, including gradient-based, input transformation-based, and model-related attacks, \etc. In this work, we find that several tiny changes in the existing adversarial attacks can significantly affect the attack performance, \eg, the number of iterations and step size. Based on careful studies of existing adversarial attacks, we propose a bag of tricks to enhance adversarial transferability, including momentum initialization, scheduled step size, dual example, spectral-based input transformation, and several ensemble strategies. Extensive experiments on the ImageNet dataset validate the high effectiveness of our proposed tricks and show that combining them can further boost adversarial transferability. Our work provides practical insights and techniques to enhance adversarial transferability, and offers guidance to improve the attack performance on the real-world application through simple adjustments.
LightHouse: A Survey of AGI Hallucination
With the development of artificial intelligence, large-scale models have become increasingly intelligent. However, numerous studies indicate that hallucinations within these large models are a bottleneck hindering the development of AI research. In the pursuit of achieving strong artificial intelligence, a significant volume of research effort is being invested in the AGI (Artificial General Intelligence) hallucination research. Previous explorations have been conducted in researching hallucinations within LLMs (Large Language Models). As for multimodal AGI, research on hallucinations is still in an early stage. To further the progress of research in the domain of hallucinatory phenomena, we present a bird's eye view of hallucinations in AGI, summarizing the current work on AGI hallucinations and proposing some directions for future research.