Goto

Collaborating Authors

 Government


On the Complexity of the Bipartite Polarization Problem: from Neutral to Highly Polarized Discussions

arXiv.org Artificial Intelligence

The Bipartite Polarization Problem is an optimization problem where the goal is to find the highest polarized bipartition on a weighted and labelled graph that represents a debate developed through some social network, where nodes represent user's opinions and edges agreement or disagreement between users. This problem can be seen as a generalization of the maxcut problem, and in previous work approximate solutions and exact solutions have been obtained for real instances obtained from Reddit discussions, showing that such real instances seem to be very easy to solve. In this paper, we investigate further the complexity of this problem, by introducing an instance generation model where a single parameter controls the polarization of the instances in such a way that this correlates with the average complexity to solve those instances. The average complexity results we obtain are consistent with our hypothesis: the higher the polarization of the instance, the easier is to find the corresponding polarized bipartition.


Model Reporting for Certifiable AI: A Proposal from Merging EU Regulation into AI Development

arXiv.org Artificial Intelligence

Despite large progress in Explainable and Safe AI, practitioners suffer from a lack of regulation and standards for AI safety. In this work we merge recent regulation efforts by the European Union and first proposals for AI guidelines with recent trends in research: data and model cards. We propose the use of standardized cards to document AI applications throughout the development process. Our main contribution is the introduction of use-case and operation cards, along with updates for data and model cards to cope with regulatory requirements. We reference both recent research as well as the source of the regulation in our cards and provide references to additional support material and toolboxes whenever possible. The goal is to design cards that help practitioners develop safe AI systems throughout the development process, while enabling efficient third-party auditing of AI applications, being easy to understand, and building trust in the system. Our work incorporates insights from interviews with certification experts as well as developers and individuals working with the developed AI applications.


A data science axiology: the nature, value, and risks of data science

arXiv.org Artificial Intelligence

Data Systems Laboratory, School of Engineering and Applied Sciences Harvard University, Cambridge, MA USA =============DRAFT July 18, 2023====================== Data science is not a science. It is a research a theory of value that defines the nature, value, paradigm. As data science is in its surpass science - our most powerful research infancy, its axiology can only be speculated. Such paradigm - in enabling knowledge discovery that an axiology can aid in understanding and defining is changing our world[10]. This paper explores and data science and recognizing potenUal benefits, evaluates its remarkable, definiUve features. We present the history and nature of data science and offer Modern data science is in its infancy. Emerging candidate definiUons of essenUal data science slowly since 1962 and rapidly since 2000, data concepts required to discuss its axiology. Within a science is a fundamentally new field of inquiry, decade, this remarkable new research paradigm one of the most acUve, powerful, and rapidly will be seen as a milestone in human knowledge evolving innovaUons of the 21st century. Yet we are just beginning to data science as a Promethean Moment[10] that understand and define it. Due to based on single invenUons, e.g., the prinUng press, its infancy, many definiUons are independent, this moment is based on a meta-technology Essen'al data science concepts data science community to achieve such a Data science (the data science research paradigm) definiUon. To problem solving based on its unique ability to contribute to an iniUal assessment and definiUon computaUonally analyze data to discover insights of data science, this paper proposes an iniUal into moUvaUng domain problems where the axiology of data science. A comprehensive data science axiology is (i.e., learning from data) of data science research A meta technology is used to produce new technology and knowledge hence can be applicable to most human endeavors. Data about, discover, arUculate, and validate the true science results are probabilis5c, correla5onal, nature of the ul5mate ques5ons about natural, possibly fragile or specific to the analysis method observable phenomena as new knowledge about or dataset, cannot be proven complete or correct, those phenomena. ScienUfic results are defini5ve, and lack explana5ons and interpreta5ons for the conclusive, casual, robust, universal knowledge of mo5va5ng domain problem[46]. Like all research paradigms, science and discovery conducted by applying the data science data science are complementary.


IsoEx: an explainable unsupervised approach to process event logs cyber investigation

arXiv.org Artificial Intelligence

39 seconds. That is the timelapse between two consecutive cyber attacks as of 2023. Meaning that by the time you are done reading this abstract, about 1 or 2 additional cyber attacks would have occurred somewhere in the world. In this context of highly increased frequency of cyber threats, Security Operation Centers (SOC) and Computer Emergency Response Teams (CERT) can be overwhelmed. In order to relieve the cybersecurity teams in their investigative effort and help them focus on more added-value tasks, machine learning approaches and methods started to emerge. This paper introduces a novel method, IsoEx, for detecting anomalous and potentially problematic command lines during the investigation of contaminated devices. IsoEx is built around a set of features that leverages the log structure of the command line, as well as its parent/child relationship, to achieve a greater accuracy than traditional methods. To detect anomalies, IsoEx resorts to an unsupervised anomaly detection technique that is both highly sensitive and lightweight. A key contribution of the paper is its emphasis on interpretability, achieved through the features themselves and the application of eXplainable Artificial Intelligence (XAI) techniques and visualizations. This is critical to ensure the adoption of the method by SOC and CERT teams, as the paper argues that the current literature on machine learning for log investigation has not adequately addressed the issue of explainability. This method was proven efficient in a real-life environment as it was built to support a company\'s SOC and CERT


On Provable Copyright Protection for Generative Models

arXiv.org Artificial Intelligence

There is a growing concern that learned conditional generative models may output samples that are substantially similar to some copyrighted data $C$ that was in their training set. We give a formal definition of $\textit{near access-freeness (NAF)}$ and prove bounds on the probability that a model satisfying this definition outputs a sample similar to $C$, even if $C$ is included in its training set. Roughly speaking, a generative model $p$ is $\textit{$k$-NAF}$ if for every potentially copyrighted data $C$, the output of $p$ diverges by at most $k$-bits from the output of a model $q$ that $\textit{did not access $C$ at all}$. We also give generative model learning algorithms, which efficiently modify the original generative model learning algorithm in a black box manner, that output generative models with strong bounds on the probability of sampling protected content. Furthermore, we provide promising experiments for both language (transformers) and image (diffusion) generative models, showing minimal degradation in output quality while ensuring strong protections against sampling protected content.


NusaCrowd: Open Source Initiative for Indonesian NLP Resources

arXiv.org Artificial Intelligence

We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought together 137 datasets and 118 standardized data loaders. The quality of the datasets has been assessed manually and automatically, and their value is demonstrated through multiple experiments. NusaCrowd's data collection enables the creation of the first zero-shot benchmarks for natural language understanding and generation in Indonesian and the local languages of Indonesia. Furthermore, NusaCrowd brings the creation of the first multilingual automatic speech recognition benchmark in Indonesian and the local languages of Indonesia. Our work strives to advance natural language processing (NLP) research for languages that are under-represented despite being widely spoken.


A picture of the space of typical learnable tasks

arXiv.org Artificial Intelligence

We develop information geometric techniques to understand the representations learned by deep networks when they are trained on different tasks using supervised, meta-, semi-supervised and contrastive learning. We shed light on the following phenomena that relate to the structure of the space of tasks: (1) the manifold of probabilistic models trained on different tasks using different representation learning methods is effectively low-dimensional; (2) supervised learning on one task results in a surprising amount of progress even on seemingly dissimilar tasks; progress on other tasks is larger if the training task has diverse classes; (3) the structure of the space of tasks indicated by our analysis is consistent with parts of the Wordnet phylogenetic tree; (4) episodic meta-learning algorithms and supervised learning traverse different trajectories during training but they fit similar models eventually; (5) contrastive and semi-supervised learning methods traverse trajectories similar to those of supervised learning. We use classification tasks constructed from the CIFAR-10 and Imagenet datasets to study these phenomena.


Forecasting consumer confidence through semantic network analysis of online news

arXiv.org Artificial Intelligence

This research studies the impact of online news on social and economic consumer perceptions through semantic network analysis. Using over 1.8 million online articles on Italian media covering four years, we calculate the semantic importance of specific economic-related keywords to see if words appearing in the articles could anticipate consumers' judgments about the economic situation and the Consumer Confidence Index. We use an innovative approach to analyze big textual data, combining methods and tools of text mining and social network analysis. Results show a strong predictive power for the judgments about the current households and national situation. Our indicator offers a complementary approach to estimating consumer confidence, lessening the limitations of traditional survey-based methods.


The Man Who Wrote the AI Doomer Bible

The Atlantic - Technology

A framed photograph of three men in military fatigues hangs above his desk. They're tightening straps on what first appear to be two water heaters but are, in fact, thermonuclear weapons. Resting against a nearby wall is a black-and-white print depicting the first billionth of a second after the detonation of an atomic bomb: a thousand-foot-tall ghostly amoeba. And above us, dangling from the ceiling like the sword of Damocles, is a plastic model of the Hindenburg. Depending on how you choose to look at it, Rhodes's office is either a shrine to awe-inspiring technological progress or a harsh reminder of its power to incinerate us all in the blink of an eye.


Donald Trump needs someone to play legal traffic cop

Slate

This week, Emily Bazelon, John Dickerson, and David Plotz are together again and talking about Donald Trump's next indictment and the charges against his "false electors" in Michigan; the struggles of candidates Ron DeSantis, Tim Scott, et al.; and Congressional Republicans' culture war against the U.S. military. Here are some notes and references from this week's show: Fox News Digital: "Republican presidential candidate Sen. Tim Scott says Donald Trump is'overqualified to be my vice president'" Manu Raju, Rashard Rose, and Lauren Fox for CNN: "Tommy Tuberville now says'White nationalists are racists' after refusing to denounce them" Zoë Richards for NBC News: "Arizona Republican refers to Black Americans as'colored people' in House floor debate" Here are this week's chatters: John: Mona El-Naggar, Johan M. Kessel, and Alexander Stockton for The New York Times: "What Is War to a Grieving Child?"; Jeanna Smialek and Ben Casselman for The New York Times: "The Pandemic's Labor Market Myths"; and Chris Cameron for The New York Times: "Over 700 Civil War-Era Gold Coins Found Buried on a Kentucky Farm" David: "Exploring a Secret Fort" with David through airbnb; Steve Bohnel for The Frederick News-Post: "$200,000, or the city burns: The story of the Confederacy's ransom on Frederick"; and Caity Weaver for The New York Times Magazine: "My Impossible Mission to Find Tom Cruise" For this week's Slate Plus bonus segment, David, John, and Emily discuss the Hollywood actors' and writers' strikes, artificial intelligence, and the future of work.