Government
A Link to News Site Meduza Can (Technically) Land You in Russian Prison
When you run a major app, all it takes is one mistake to put countless people at risk. Such is the case with Diksha, a public education app run by India's Ministry of Education that exposed the personal information of around 1 million teachers and millions of students across the country. The data, which included things like full names, email addresses, and phone numbers, was publicly accessible for at least a year and likely longer, potentially exposing those impacted to phishing attacks and other scams. Speaking of cybercrime, the LockBit ransomware gang has long operated under the radar, thanks to its professional operation and choice of targets. But over the past year, a series of missteps and drama have thrust it into the spotlight, potentially threatening its ability to continue operating with impunity.
How learners produce data from text in classifying clickbait
Horton, Nicholas J., Chao, Jie, Palmer, Phebe, Finzer, William
Text provides a compelling example of unstructured data that can be used to motivate and explore classification problems. Challenges arise regarding the representation of features of text and student linkage between text representations as character strings and identification of features that embed connections with underlying phenomena. In order to observe how students reason with text data in scenarios designed to elicit certain aspects of the domain, we employed a task-based interview method using a structured protocol with six pairs of undergraduate students. Our goal was to shed light on students' understanding of text as data using a motivating task to classify headlines as "clickbait" or "news". Three types of features (function, content, and form) surfaced, the majority from the first scenario. Our analysis of the interviews indicates that this sequence of activities engaged the participants in thinking at both the human-perception level and the computer-extraction level and conceptualizing connections between them.
Large Language Models as Corporate Lobbyists
We demonstrate a proof-of-concept of a large language model conducting corporate lobbying related activities. An autoregressive large language model (OpenAI's text-davinci-003) determines if proposed U.S. Congressional bills are relevant to specific public companies and provides explanations and confidence levels. For the bills the model deems as relevant, the model drafts a letter to the sponsor of the bill in an attempt to persuade the congressperson to make changes to the proposed legislation. We use hundreds of novel ground-truth labels of the relevance of a bill to a company to benchmark the performance of the model. It outperforms the baseline of predicting the most common outcome of irrelevance. We also benchmark the performance of the previous OpenAI GPT-3 model (text-davinci-002), which was the state-of-the-art model on many academic natural language tasks until text-davinci-003 was recently released. The performance of text-davinci-002 is worse than the simple baseline. Longer-term, if AI begins to influence law in a manner that is not a direct extension of human intentions, this threatens the critical role that law as information could play in aligning AI with humans. Initially, AI is being used to simply augment human lobbyists for a small portion of their daily tasks. However, firms have an incentive to use less and less human oversight over automated assessments of policy ideas and the written communication to regulatory agencies and Congressional staffers. The core question raised is where to draw the line between human-driven and AI-driven policy influence.
Semantic Parsing for Conversational Question Answering over Knowledge Graphs
Perez-Beltrachini, Laura, Jain, Parag, Monti, Emilio, Lapata, Mirella
In this paper, we are interested in developing semantic parsers which understand natural language questions embedded in a conversation with a user and ground them to formal queries over definitions in a general purpose knowledge graph (KG) with very large vocabularies (covering thousands of concept names and relations, and millions of entities). To this end, we develop a dataset where user questions are annotated with Sparql parses and system answers correspond to execution results thereof. We present two different semantic parsing approaches and highlight the challenges of the task: dealing with large vocabularies, modelling conversation context, predicting queries with multiple entities, and generalising to new questions at test time. We hope our dataset will serve as useful testbed for the development of conversational semantic parsers. Our dataset and models are released at https://github.com/EdinburghNLP/SPICE.
Informational Diversity and Affinity Bias in Team Growth Dynamics
Heidari, Hoda, Barocas, Solon, Kleinberg, Jon, Levy, Karen
Prior work has provided strong evidence that, within organizational settings, teams that bring a diversity of information and perspectives to a task are more effective than teams that do not. If this form of informational diversity confers performance advantages, why do we often see largely homogeneous teams in practice? One canonical argument is that the benefits of informational diversity are in tension with affinity bias. To better understand the impact of this tension on the makeup of teams, we analyze a sequential model of team formation in which individuals care about their team's performance (captured in terms of accurately predicting some future outcome based on a set of features) but experience a cost as a result of interacting with teammates who use different approaches to the prediction task. Our analysis of this simple model reveals a set of subtle behaviors that team-growth dynamics can exhibit: (i) from certain initial team compositions, they can make progress toward better performance but then get stuck partway to optimally diverse teams; while (ii) from other initial compositions, they can also move away from this optimal balance as the majority group tries to crowd out the opinions of the minority. The initial composition of the team can determine whether the dynamics will move toward or away from performance optimality, painting a path-dependent picture of inefficiencies in team compositions. Our results formalize a fundamental limitation of utility-based motivations to drive informational diversity in organizations and hint at interventions that may improve informational diversity and performance simultaneously.
Generating Random SAT Instances: Multiple Solutions could be Predefined and Deeply Hidden
Zhao, Dongdong (Wuhan University of Technology) | Liao, Lei | Luo, Wenjian | Xiang, Jianwen | Jiang, Hao | Hu, Xiaoyi
The generation of SAT instances is an important issue in computer science, and it is useful for researchers to verify the effectiveness of SAT solvers. Addressing this issue could inspire researchers to propose new search strategies. SAT problems exist in various real-world applications, some of which have more than one solution. However, although several algorithms for generating random SAT instances have been proposed, few can be used to generate hard instances that have multiple predefined solutions. In this paper, we propose the KHidden-M algorithm to generate SAT instances with multiple predefined solutions that could be hard to solve by the local search strategy when the number of predefined solutions is small enough and the Hamming distance between them is not less than half of the solution length. Specifically, first, we generate an SAT instance that is satisfied by all of the predefined solutions. Next, if the generated SAT instance does not satisfy the hardness condition, then a strategy will be conducted to adjust clauses through multiple iterations to improve the hardness of the whole instance. We propose three strategies to generate the SAT instance in the first part. The first strategy is called the random strategy, which randomly generates clauses that are satisfied by all of the predefined solutions. The other two strategies are called the estimating strategy and greedy strategy, and using them, we attempt to generate an instance that directly satisfies or is closer to the hardness condition for the local search strategy. We employ two SAT solvers (i.e., WalkSAT and Kissat) to investigate the hardness of the SAT instances generated by our algorithm in the experiments. The experimental results show the effectiveness of the random, estimating and greedy strategies. Compared to the state-of-the-art algorithm for generating SAT instances with predefined solutions, namely, M-hidden, our algorithm could be more effective in generating hard SAT instances.
Russian shelling leaves 10 Ukrainian civilians dead, 20 injured, Zelenskyy says
Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. A new barrage of Russian shelling killed at least 10 Ukrainian civilians and wounded 20 others in a day, the office of Ukraine's president said Friday as the country worked to recover from an earlier wave of Russian missile strikes and drone attacks. Regional officials said towns and villages in the east and in the south that are within reach of the Russian artillery suffered most. Six people died in the Donetsk region, two in Kherson, and two in the Kharkiv region.
The Morning After: Will AI be your next lawyer?
In a new study, University of Minnesota law professors used ChatGPT AI chatbot to answer graduate exams at four courses in their school. The AI passed all four, but with an average grade of C . The University of Minnesota group noted ChatGPT was good at addressing "basic legal rules" and summaries, but it floundered when trying to pinpoint issues relevant in a case. When faced with business management questions in a different study, the generator was "amazing" with simple operations management and process analysis questions, but it couldn't handle advanced process questions. It even made mistakes with sixth-grade-level math โ something other AI authors have struggled with.
NIST releases framework to boost risk-free adoption of AI
National Institute of Standards and Technology (NIST), a US-based federal agency responsible for building technology standards, has released artificial intelligence risk management framework (AI RMF 1.0), which can be used by companies to build and use AI systems in an ethical and risk-free manner. Developed in collaboration with private and public sector organisations, AI RMF framework is voluntary, which means it's usage is not binding on any company. However, NIST director Laurie E. Locascio believes that it can help large and small organisations across sectors manage their AI related risks more effectively. The framework is part of NIST's larger goal of "cultivating trust" in AI technologies within all communities, added Locascio. "It should accelerate AI innovation and growth while advancing -- rather than restricting or damaging -- civil rights, civil liberties and equity for all," Don Graves, Deputy Commerce Secretary, said in a statement.
Japan tightens Russia sanctions, expands export ban list
Japan has tightened its sanctions against Russia following its latest wave of missile attacks in Ukraine, adding goods to an export ban list and freezing the assets of Russian officials and entities. The decision on Friday comes after Russia launched missile attacks across Ukraine on Thursday, killing at least 11 people, following a pledge by Germany and the United States to supply tanks that could help Kyiv counter a new Russian offensive. "In light of the situation surrounding Ukraine and to contribute to international efforts to secure peace, Japan will implement export bans in line with other major nations," Japan's Ministry of Economy Trade and Industry said in a press release. Among the new sanctions, Japan will prohibit shipments of items to 49 organisations in Russia from February 3 that could be used to enhance Moscow's military capability. Those will include products ranging from water cannons, gas exploration equipment and semiconductor equipment to vaccines, X-ray inspection equipment, explosives and robots, the ministry said.