Goto

Collaborating Authors

 Government


Polynomial time algorithms for dual volume sampling

Neural Information Processing Systems

We study dual volume sampling, a method for selecting k columns from an n m short and wide matrix (n apple k apple m) such that the probability of selection is proportional to the volume spanned by the rows of the induced submatrix. This method was proposed by Avron and Boutsidis (2013), who showed it to be a promising method for column subset selection and its multiple applications. However, its wider adoption has been hampered by the lack of polynomial time sampling algorithms. We remove this hindrance by developing an exact (randomized) polynomial time sampling algorithm as well as its derandomization. Thereafter, we study dual volume sampling via the theory of real stable polynomials and prove that its distribution satisfies the "Strong Rayleigh" property. This result has numerous consequences, including a provably fast-mixing Markov chain sampler that makes dual volume sampling much more attractive to practitioners. This sampler is closely related to classical algorithms for popular experimental design methods that are to date lacking theoretical analysis but are known to empirically work well.


California Gov. Newsom signs bills to protect children from AI deepfake nudes

FOX News

Victims of A.I. deepfake pornography Elliston Berry and Francesca Mani speak out about how they were victims of fake images circulated online and how this impacted their personal lives, including at school. California Gov. Gavin Newsom signed two bills on Sunday to help protect minors from harmful sexual imagery of children created through the misuse of artificial intelligence tools. Supporters of the bills say that current law does not allow district attorneys to prosecute those who possess or distribute AI-generated child sexual abuse images if they cannot prove the materials are depicting a real person. Under the new laws, such an offense would qualify as a felony. Last month, Newsom signed legislation regulating AI-generated "deepfake" election content and requiring the removal of "deceptive content" from social media.


FALKON: An Optimal Large Scale Kernel Method

Neural Information Processing Systems

Kernel methods provide a principled way to perform non linear, nonparametric learning. They rely on solid functional analytic foundations and enjoy optimal statistical properties. However, at least in their basic form, they have limited applicability in large scale scenarios because of stringent computational requirements in terms of time and especially memory. In this paper, we take a substantial step in scaling up kernel methods, proposing FALKON, a novel algorithm that allows to efficiently process millions of points. FALKON is derived combining several algorithmic principles, namely stochastic subsampling, iterative solvers and preconditioning. Our theoretical analysis shows that optimal statistical accuracy is achieved requiring essentially O(n) memory and O(n n) time. An extensive experimental analysis on large scale datasets shows that, even with a single machine, FALKON outperforms previous state of the art solutions, which exploit parallel/distributed architectures.


Fox News AI Newsletter: 'Setback' for AI safety

FOX News

Fox News chief political anchor Bret Baier has the latest on the pros and cons of the bombshell developments on'Special Report.' California Gov. Gavin Newsom speaks to reporters after a presidential debate between President Joe Biden and Republican presidential candidate former President Donald Trump in Atlanta on Thursday, June 27, 2024. SAFETY SETBACK: California Gov. Gavin Newsom, a Democrat, on Sunday vetoed a bill to create safety measures for large artificial intelligence models, which would have been the first such law in the nation. AI BOOM: Investors in the behemoth SPDR technology sector fund might be surprised to learn that until last week their exposure to Nvidia was roughly four times that of Apple, despite their comparable market values. SF'S CRIME WEAPON: San Francisco is taking a bold step in its fight against crime by deploying three new mobile surveillance cameras.


Shh, ChatGPT. That's a Secret.

The Atlantic - Technology

This past spring, a man in Washington State worried that his marriage was on the verge of collapse. "I am depressed and going a little crazy, still love her and want to win her back," he typed into ChatGPT. With the chatbot's help, he wanted to write a letter protesting her decision to file for divorce and post it to their bedroom door. "Emphasize my deep guilt, shame, and remorse for not nurturing and being a better husband, father, and provider," he wrote. In another message, he asked ChatGPT to write his wife a poem "so epic that it could make her change her mind but not cheesy or over the top." The man's chat history was included in the WildChat data set, a collection of 1 million ChatGPT conversations gathered consensually by researchers to document how people are interacting with the popular chatbot.


The Solution to Giant Killer Cars Is Really, Really Simple

Slate

The number of pedestrian deaths in the United States is skyrocketing. In 2022 traffic crashes killed 7,805 people on foot--that's an 83 percent rise from 2009, and a 40-year high. The vast majority of those deaths involved a car colliding into a human. In September, the National Highway Traffic Safety Administration took a significant step toward addressing the crisis of killer cars, proposing a new federal rule that would require--not ask--carmakers to ensure that the front ends of their vehicles do not create excessive risk of pedestrian head injuries. Should the proposal become law, hulking SUVs and pickups would face particular challenges passing NHTSA's mandatory tests.


The Good Robot podcast: the EU AI Act part 2, with Amba Kak and Sarah Myers West from AI NOW

AIHub

Hosted by Eleanor Drage and Kerry Mackereth, The Good Robot is a podcast which explores the many complex intersections between gender, feminism and technology. In the second instalment of our EU AI Act series we talk to Amba Kak and Sarah Myers West, the Co-Directors of the AI Now Institute, a leading policy thinktank based in New York. Amba and Sarah talk about why policy narratives matter, why it's actually fake news that AI is moving too fast for regulation to follow, and why innovation versus regulation is a lazy and outdated maxim. Meanwhile, we chip in with some weird comments about why kitchen whisks are awesome, and why getting inundated by emails is the present day equivalent of somebody badgering your cows in the 1800s. Don't forget to check out our first instalment of the EU AI Act series with Daniel Leufer and Caterina Daniels from Access Now, which is available on YouTube, Spotify, Apple, or any of your other favourite podcasting platforms.


Eight Scientists, a Billion Dollars, and the Moonshot Agency Trying to Make Britain Great Again

WIRED

In a cramped conference room in Bristol, Ilan Gur is trying to convince a group of plant biologists that they can change the world. The 44-year-old has the patter you'd expect from a Californian startup founder, but he's also one of the UK's most senior civil servants, so what comes next is unexpected. Close your eyes, he asks the scientists, and imagine pushing past the very edges of your research. The attendees take a beat, shifting slightly on their uncomfortable chairs. Positive visualization is not quite what they had expected from a workshop introducing them to the Advanced Research and Invention Agency (ARIA), the UK government's new high-risk, high-reward science funding agency.


CALF: Benchmarking Evaluation of LFQA Using Chinese Examinations

arXiv.org Artificial Intelligence

Long-Form Question Answering (LFQA) refers to generating in-depth, paragraph-level responses to open-ended questions. Although lots of LFQA methods are developed, evaluating LFQA effectively and efficiently remains challenging due to its high complexity and cost. Therefore, there is no standard benchmark for LFQA evaluation till now. To address this gap, we make the first attempt by proposing a well-constructed, reference-based benchmark named Chinese exAmination for LFQA Evaluation (CALF), aiming to rigorously assess the performance of automatic evaluation metrics for LFQA. The CALF benchmark is derived from Chinese examination questions that have been translated into English. It includes up to 1476 examples consisting of knowledge-intensive and nuanced responses. Our evaluation comprises three different settings to ana lyze the behavior of automatic metrics comprehensively. We conducted extensive experiments on 7 traditional evaluation metrics, 3 prompt-based metrics, and 3 trained evaluation metrics, and tested on agent systems for the LFQA evaluation. The results reveal that none of the current automatic evaluation metrics shows comparable performances with humans, indicating that they cannot capture dense information contained in long-form responses well. In addition, we provide a detailed analysis of the reasons why automatic evaluation metrics fail when evaluating LFQA, offering valuable insights to advance LFQA evaluation systems. Dataset and associated codes can be accessed at our GitHub repository.


Real-World Data and Calibrated Simulation Suite for Offline Training of Reinforcement Learning Agents to Optimize Energy and Emission in Buildings for Environmental Sustainability

arXiv.org Artificial Intelligence

Commercial office buildings contribute 17 percent of Carbon Emissions in the US, according to the US Energy Information Administration (EIA), and improving their efficiency will reduce their environmental burden and operating cost. A major contributor of energy consumption in these buildings are the Heating, Ventilation, and Air Conditioning (HVAC) devices. HVAC devices form a complex and interconnected thermodynamic system with the building and outside weather conditions, and current setpoint control policies are not fully optimized for minimizing energy use and carbon emission. Given a suitable training environment, a Reinforcement Learning (RL) agent is able to improve upon these policies, but training such a model, especially in a way that scales to thousands of buildings, presents many practical challenges. Most existing work on applying RL to this important task either makes use of proprietary data, or focuses on expensive and proprietary simulations that may not be grounded in the real world. We present the Smart Buildings Control Suite, the first open source interactive HVAC control dataset extracted from live sensor measurements of devices in real office buildings. The dataset consists of two components: six years of real-world historical data from three buildings, for offline RL, and a lightweight interactive simulator for each of these buildings, calibrated using the historical data, for online and model-based RL. For ease of use, our RL environments are all compatible with the OpenAI gym environment standard. We also demonstrate a novel method of calibrating the simulator, as well as baseline results on training an RL agent on the simulator, predicting real-world data, and training an RL agent directly from data. We believe this benchmark will accelerate progress and collaboration on building optimization and environmental sustainability research.