Government
Large Language Models Can Be Strong Differentially Private Learners
Li, Xuechen, Tramèr, Florian, Liang, Percy, Hashimoto, Tatsunori
Differentially Private (DP) learning has seen limited success for building large deep learning models of text, and straightforward attempts at applying Differentially Private Stochastic Gradient Descent (DP-SGD) to NLP tasks have resulted in large performance drops and high computational overhead. We show that this performance drop can be mitigated with (1) the use of large pretrained language models; (2) non-standard hyperparameters that suit DP optimization; and (3) fine-tuning objectives which are aligned with the pretraining procedure. With the above, we obtain NLP models that outperform state-of-the-art DP-trained models under the same privacy budget and strong non-private baselines -- by directly fine-tuning pretrained models with DP optimization on moderately-sized corpora. To address the computational challenge of running DP-SGD with large Transformers, we propose a memory saving technique that allows clipping in DP-SGD to run without instantiating per-example gradients for any linear layer in the model. The technique enables privately training Transformers with almost the same memory cost as non-private training at a modest run-time overhead. Contrary to conventional wisdom that DP optimization fails at learning high-dimensional models (due to noise that scales with dimension) empirical results reveal that private learning with pretrained language models doesn't tend to suffer from dimension-dependent performance degradation. Code to reproduce results can be found at https://github.com/lxuechen/private-transformers.
DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question Answering
Neeman, Ella, Aharoni, Roee, Honovich, Or, Choshen, Leshem, Szpektor, Idan, Abend, Omri
Question answering models commonly have access to two sources of "knowledge" during inference time: (1) parametric knowledge - the factual knowledge encoded in the model weights, and (2) contextual knowledge - external knowledge (e.g., a Wikipedia passage) given to the model to generate a grounded answer. Having these two sources of knowledge entangled together is a core issue for generative QA models as it is unclear whether the answer stems from the given non-parametric knowledge or not. This unclarity has implications on issues of trust, interpretability and factuality. In this work, we propose a new paradigm in which QA models are trained to disentangle the two sources of knowledge. Using counterfactual data augmentation, we introduce a model that predicts two answers for a given question: one based on given contextual knowledge and one based on parametric knowledge. Our experiments on the Natural Questions dataset show that this approach improves the performance of QA models by making them more robust to knowledge conflicts between the two knowledge sources, while generating useful disentangled answers.
The Metaverse Data Deluge: What Can We Do About It?
Ooi, Beng Chin, Chen, Gang, Shou, Mike Zheng, Tan, Kian-Lee, Tung, Anthony, Xiao, Xiaokui, Yip, James Wei Luen, Zhang, Meihui
In the Metaverse, the physical space and the virtual space co-exist, and interact simultaneously. While the physical space is virtually enhanced with information, the virtual space is continuously refreshed with real-time, real-world information. To allow users to process and manipulate information seamlessly between the real and digital spaces, novel technologies must be developed. These include smart interfaces, new augmented realities, efficient storage and data management and dissemination techniques. In this paper, we first discuss some promising co-space applications. These applications offer opportunities that neither of the spaces can realize on its own. We then discuss challenges. Finally, we discuss and envision what are likely to be required from the database and system perspectives.
Artificial Intelligence (AI) Takes a Role in USPTO Patent Searches
In 2021 the U.S. Patent and Trademark Office (USPTO) developed an Artificial Intelligence (AI) based prototype search system for use by examiners during examination of patent applications. As previously discussed by Mintz, the AI search system aimed to help identify relevant documents and provide suggestions to examiners for additional areas to search. The USPTO found searching success with the prototype, for the USPTO just launched an AI-based "Similarity Search" in the Patents End-to-End (PE2E) prior art search suite for patents examiners. As explained by the USPTO, a patent examiner provides input, including a patent specification, to the "Similarity Search" feature. The feature then uses AI models to identify and, within seconds, output U.S. and foreign patent references similar to the patent application being examined.
Iran Calls For Ukraine Talks As It Hosts Russian Security Chief
Iran's top security official Ali Shamkhani called for dialogue to end the war in Ukraine during a meeting Wednesday in Tehran with his Russian counterpart Nikolai Patrushev. Patrushev later met with President Ebrahim Raisi, who said that "broadening the scope of the war and its escalation are a source of concern for all countries," according to official news agency IRNA. The meetings come after Kiev and its Western allies accused Russia in recent weeks of using Iranian-made drones to carry out attacks in Ukraine. "Iran supports any initiative leading to a ceasefire and peace between Russia and Ukraine based on dialogue," Shamkhani said, secretary of Iran's Supreme National Security Council. The Islamic republic was "ready to play a role in ending the war", he added, IRNA reported.
How AI Can Help Government Deliver Better Outcomes
Transcript: How can AI positively transform the way we work, especially in the public sector? Over the past year and a half, we've seen a lot of different ways public servants are using AI to help them meet their mission and transform the constituent experience. AI can surface information insights more rapidly and at scale, and what's really interesting is that it can help agencies translate, "gov speak" to "people speak." AI can build that connective tissue between how people speak and how they actually live their lives and look for information, versus how government has organized it. Also, our public servants are burnt out.
Artificial Intelligence (AI) Takes a Role in USPTO Patent Searches
In 2021 the U.S. Patent and Trademark Office (USPTO) developed an Artificial Intelligence (AI) based prototype search system for use by examiners during examination of patent applications. As previously discussed by Mintz, the AI search system aimed to help identify relevant documents and provide suggestions to examiners for additional areas to search. The USPTO found searching success with the prototype, for the USPTO just launched an AI-based "Similarity Search" in the Patents End-to-End (PE2E) prior art search suite for patents examiners. As explained by the USPTO, a patent examiner provides input, including a patent specification, to the "Similarity Search" feature. The feature then uses AI models to identify and, within seconds, output U.S. and foreign patent references similar to the patent application being examined.
The Download: capturing carbon with seagrass, and China's election interference
For years, Tidal, a project within Alphabet's "moonshot factory" X division, has been using cameras, computer vision and machine learning to get a better understanding of life beneath the oceans, including monitoring fish off the coast of Norway. Now, MIT Technology Review can report, Tidal hopes its system can help preserve and restore the world's seagrass beds, accelerating efforts to harness the oceans to suck up and store away far more carbon dioxide. The project's ambitious mission is to improve our understanding of underwater ecosystems in order to inform and incentivize efforts to protect the oceans amid mounting threats. It could also provide crucial answers to the many questions hanging over seagrass' role in both sucking up carbon and regulating the climate. China is copying Russia's election interference playbook China is increasingly interfering in US politics by getting its agents to create social media accounts posing as American citizens, according to research co-led by Renée DiResta, the technical research manager at the Stanford Internet Observatory, who has studied foreign influence on social media for years.
Robot to make history by speaking in House of Lords debate
The world's first ultra-realistic robot artist is set to make history as the first robot to speak at the House of Lords as she questions whether creativity is under attack from the rise of Artificial Intelligence and technology. Ai-Da will give a fresh perspective on the role of technology in creating art in the future, how AI art differs to what human artists produce, and the limits of technology in creating art. It will argue that while machine creativity presents a great opporunity for us to explore new ideas and ways of thinking, "there are also risks associated with this technology which we need to consider carefully". "We need to think of benefits and limitations, and consider ethical implications." Director of the Ai-Da Robot project, Aidan Meller comments: "Creativity is changing. Ai-Da challenges what it means to be an artist in a post-human world. "Her abilities as an artist brings into question the foundations of the art world and the creative industries.