Government
Google fires 28 staff after protests against cloud contract with Israel
Google has fired 28 employees following a sit-down protest over the tech giant's contract to provide cloud computing and artificial intelligence services to the Israeli government The terminations come after the group No Tech for Apartheid on Tuesday occupied Google offices in California and New York to protest the 1.2bn contract known as Project Nimbus. Video of the demonstrations shared on social media showed police arresting employees in the office of Google Cloud CEO Thomas Kurian. In a statement on Thursday, Google said that physically impeding employees and preventing them from accessing company facilities was a "clear violation of our policies and completely unacceptable behaviour". "After refusing multiple requests to leave the premises, law enforcement was engaged to remove them to ensure office safety," a spokesperson said. "We have so far concluded individual investigations that resulted in the termination of employment for 28 employees, and will continue to investigate and take action as needed."
Israeli missiles hit site in Iran, media report says
Israeli missiles have hit a site in Iran, ABC News reported late on Thursday, citing a U.S. official, days after Iran launched a drone strike on Israel in response to an attack at the Iranian embassy in Syria. Iran's Fars news agency said an explosion was heard at an airport in the Iranian city of Isafahan, but the cause was not immediately known. Several Iranian nuclear sites are located in Isfahan province, including Natanz, the centerpiece of Iran's uranium enrichment program. Several flights were diverted over Iranian airspace, CNN reported. Over the weekend, Iran launched hundreds of drones and missiles in a retaliatory strike after a suspected Israeli strike on its embassy compound in Syria.
A drone strike in Odesa shatters a family's life
In the photograph, Anna Haidarzhy and her 4-month-old son, Tymofii, are barely visible under the bloodstained blanket. They lie in the rubble, at the feet of rescue workers in black and fluorescent uniforms. Just two arms, one from the mother, 31, one from her son, can be seen sticking out of the blanket. "It looked like they were saying goodbye," one of the rescuers, Serhii Mudrenko, said of the image. Their bodies were found in the smoking ruins of an apartment block hit in a Russian drone attack in March in the southern Ukrainian city of Odesa that killed 12 people.
What Generative Artificial Intelligence Means for Terminological Definitions
This paper examines the impact of Generative Artificial Intelligence (GenAI) tools like ChatGPT on the creation and consumption of terminological definitions. From the terminologist's point of view, the strategic use of GenAI tools can streamline the process of crafting definitions, reducing both time and effort, while potentially enhancing quality. GenAI tools enable AI-assisted terminography, notably post-editing terminography, where the machine produces a definition that the terminologist then corrects or refines. However, the potential of GenAI tools to fulfill all the terminological needs of a user, including term definitions, challenges the very existence of terminological definitions and resources as we know them. Unlike terminological definitions, GenAI tools can describe the knowledge activated by a term in a specific context. However, a main drawback of these tools is that their output can contain errors. For this reason, users requiring reliability will likely still resort to terminological resources for definitions. Nevertheless, with the inevitable integration of AI into terminology work, the distinction between human-created and AI-created content will become increasingly blurred.
A national longitudinal dataset of skills taught in U.S. higher education curricula
Sabet, Alireza Javadian, Bana, Sarah H., Yu, Renzhe, Frank, Morgan R.
Higher education plays a critical role in driving an innovative economy by equipping students with knowledge and skills demanded by the workforce. While researchers and practitioners have developed data systems to track detailed occupational skills, such as those established by the U.S. Department of Labor (DOL), much less effort has been made to document skill development in higher education at a similar granularity. Here, we fill this gap by presenting a longitudinal dataset of skills inferred from over three million course syllabi taught at nearly three thousand U.S. higher education institutions. To construct this dataset, we apply natural language processing to extract from course descriptions detailed workplace activities (DWAs) used by the DOL to describe occupations. We then aggregate these DWAs to create skill profiles for institutions and academic majors. Our dataset offers a large-scale representation of college-educated workers and their role in the economy. To showcase the utility of this dataset, we use it to 1) compare the similarity of skills taught and skills in the workforce according to the US Bureau of Labor Statistics, 2) estimate gender differences in acquired skills based on enrollment data, 3) depict temporal trends in the skills taught in social science curricula, and 4) connect college majors' skill distinctiveness to salary differences of graduates. Overall, this dataset can enable new research on the source of skills in the context of workforce development and provide actionable insights for shaping the future of higher education to meet evolving labor demands especially in the face of new technologies.
Enhancing Q&A with Domain-Specific Fine-Tuning and Iterative Reasoning: A Comparative Study
Nguyen, Zooey, Annunziata, Anthony, Luong, Vinh, Dinh, Sang, Le, Quynh, Ha, Anh Hai, Le, Chanh, Phan, Hong An, Raghavan, Shruti, Nguyen, Christopher
AI-powered question-answering (Q&A) systems have emerged as important tools, alongside established search technologies, to enable quick access to relevant information and knowledge from large digital sources that are complex and time-consuming for humans to navigate. Advancements in large language models (LLMs) have revolutionized the field of Q&A, with models like GPT-3 (Brown et al. 2020), BERT (Devlin et al. 2018), and RoBERTa (Liu et al. 2019) demonstrating remarkable abilities in understanding and generating human-like text. However, the effectiveness of such models in handling domain-specific questions that require specialized knowledge is limited. Retrieval-augmented generation (RAG) techniques, which combine information retrieval and generative models (Lewis et al. 2021), have shown promise in boosting the quality of LLM output in Q&A tasks. RAG systems leverage the strengths of both retrieval and generation components to provide contextually relevant and informative responses. While there is a lack of established quantification of RAG accuracy, early findings suggest that generic RAG does not perform well in complex domains such as finance. In one instance, RAG based on generic LLMs such as GPT-4-Turbo fails to answer 81% of the questions derived from Securities and Exchange Commission (SEC) financial filings (Islam et al. 2023). Aitomatic, Inc. (except as noted, all authors are from Aitomatic)
Scalable Data Assimilation with Message Passing
Key, Oscar, Takao, So, Giles, Daniel, Deisenroth, Marc Peter
Data assimilation is a core component of numerical weather prediction systems. The large quantity of data processed during assimilation requires the computation to be distributed across increasingly many compute nodes, yet existing approaches suffer from synchronisation overhead in this setting. In this paper, we exploit the formulation of data assimilation as a Bayesian inference problem and apply a message-passing algorithm to solve the spatial inference problem. Since message passing is inherently based on local computations, this approach lends itself to parallel and distributed computation. In combination with a GPU-accelerated implementation, we can scale the algorithm to very large grid sizes while retaining good accuracy and compute and memory requirements.
Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them?
Longpre, Shayne, Mahari, Robert, Obeng-Marnu, Naana, Brannon, William, South, Tobin, Gero, Katy, Pentland, Sandy, Kabbara, Jad
New capabilities in foundation models are owed in large part to massive, widely-sourced, and under-documented training data collections. Existing practices in data collection have led to challenges in documenting data transparency, tracing authenticity, verifying consent, privacy, representation, bias, copyright infringement, and the overall development of ethical and trustworthy foundation models. In response, regulation is emphasizing the need for training data transparency to understand foundation models' limitations. Based on a large-scale analysis of the foundation model training data landscape and existing solutions, we identify the missing infrastructure to facilitate responsible foundation model development practices. We examine the current shortcomings of common tools for tracing data authenticity, consent, and documentation, and outline how policymakers, developers, and data creators can facilitate responsible foundation model development by adopting universal data provenance standards.
SOPHON: Non-Fine-Tunable Learning to Restrain Task Transferability For Pre-trained Models
Deng, Jiangyi, Pang, Shengyuan, Chen, Yanjiao, Xia, Liangming, Bai, Yijie, Weng, Haiqin, Xu, Wenyuan
Instead of building deep learning models from scratch, developers are more and more relying on adapting pre-trained models to their customized tasks. However, powerful pre-trained models may be misused for unethical or illegal tasks, e.g., privacy inference and unsafe content generation. In this paper, we introduce a pioneering learning paradigm, non-fine-tunable learning, which prevents the pre-trained model from being fine-tuned to indecent tasks while preserving its performance on the original task. To fulfill this goal, we propose SOPHON, a protection framework that reinforces a given pre-trained model to be resistant to being fine-tuned in pre-defined restricted domains. Nonetheless, this is challenging due to a diversity of complicated fine-tuning strategies that may be adopted by adversaries. Inspired by model-agnostic meta-learning, we overcome this difficulty by designing sophisticated fine-tuning simulation and fine-tuning evaluation algorithms. In addition, we carefully design the optimization process to entrap the pre-trained model within a hard-to-escape local optimum regarding restricted domains. We have conducted extensive experiments on two deep learning modes (classification and generation), seven restricted domains, and six model architectures to verify the effectiveness of SOPHON. Experiment results verify that fine-tuning SOPHON-protected models incurs an overhead comparable to or even greater than training from scratch. Furthermore, we confirm the robustness of SOPHON to three fine-tuning methods, five optimizers, various learning rates and batch sizes. SOPHON may help boost further investigations into safe and responsible AI.
Mapping Social Choice Theory to RLHF
Recent work on the limitations of using reinforcement learning from human feedback (RLHF) to incorporate human preferences into model behavior often raises social choice theory as a reference point. Social choice theory's analysis of settings such as voting mechanisms provides technical infrastructure that can inform how to aggregate human preferences amid disagreement. We analyze the problem settings of social choice and RLHF, identify key differences between them, and discuss how these differences may affect the RLHF interpretation of well-known technical results in social choice.