Government
Explainable Clustering via Exemplars: Complexity and Efficient Approximation Algorithms
Davidson, Ian, Livanos, Michael, Gourru, Antoine, Walker, Peter, Velcin, Julien, Ravi, S. S.
Explainable AI (XAI) is an important developing area but remains relatively understudied for clustering. We propose an explainable-by-design clustering approach that not only finds clusters but also exemplars to explain each cluster. The use of exemplars for understanding is supported by the exemplar-based school of concept definition in psychology. We show that finding a small set of exemplars to explain even a single cluster is computationally intractable; hence, the overall problem is challenging. We develop an approximation algorithm that provides provable performance guarantees with respect to clustering quality as well as the number of exemplars used. This basic algorithm explains all the instances in every cluster whilst another approximation algorithm uses a bounded number of exemplars to allow simpler explanations and provably covers a large fraction of all the instances. Experimental results show that our work is useful in domains involving difficult to understand deep embeddings of images and text.
Signal Decomposition Using Masked Proximal Operators
Meyers, Bennet E., Boyd, Stephen P.
We consider the well-studied problem of decomposing a vector time series signal into components with different characteristics, such as smooth, periodic, nonnegative, or sparse. We describe a simple and general framework in which the components are defined by loss functions (which include constraints), and the signal decomposition is carried out by minimizing the sum of losses of the components (subject to the constraints). When each loss function is the negative log-likelihood of a density for the signal component, this framework coincides with maximum a posteriori probability (MAP) estimation; but it also includes many other interesting cases. Summarizing and clarifying prior results, we give two distributed optimization methods for computing the decomposition, which find the optimal decomposition when the component class loss functions are convex, and are good heuristics when they are not. Both methods require only the masked proximal operator of each of the component loss functions, a generalization of the well-known proximal operator that handles missing entries in its argument. Both methods are distributed, i.e., handle each component separately. We derive tractable methods for evaluating the masked proximal operators of some loss functions that, to our knowledge, have not appeared in the literature.
Twitter Topic Classification
Antypas, Dimosthenis, Ushio, Asahi, Camacho-Collados, Jose, Neves, Leonardo, Silva, Vรญtor, Barbieri, Francesco
Social media platforms host discussions about a wide variety of topics that arise everyday. Making sense of all the content and organising it into categories is an arduous task. A common way to deal with this issue is relying on topic modeling, but topics discovered using this technique are difficult to interpret and can differ from corpus to corpus. In this paper, we present a new task based on tweet topic classification and release two associated datasets. Given a wide range of topics covering the most important discussion points in social media, we provide training and testing data from recent time periods that can be used to evaluate tweet classification models. Moreover, we perform a quantitative evaluation and analysis of current general- and domain-specific language models on the task, which provide more insights on the challenges and nature of the task.
Metadata Archaeology: Unearthing Data Subsets by Leveraging Training Dynamics
Siddiqui, Shoaib Ahmed, Rajkumar, Nitarshan, Maharaj, Tegan, Krueger, David, Hooker, Sara
Modern machine learning research relies on relatively few carefully curated datasets. Even in these datasets, and typically in `untidy' or raw data, practitioners are faced with significant issues of data quality and diversity which can be prohibitively labor intensive to address. Existing methods for dealing with these challenges tend to make strong assumptions about the particular issues at play, and often require a priori knowledge or metadata such as domain labels. Our work is orthogonal to these methods: we instead focus on providing a unified and efficient framework for Metadata Archaeology -- uncovering and inferring metadata of examples in a dataset. We curate different subsets of data that might exist in a dataset (e.g. mislabeled, atypical, or out-of-distribution examples) using simple transformations, and leverage differences in learning dynamics between these probe suites to infer metadata of interest. Our method is on par with far more sophisticated mitigation methods across different tasks: identifying and correcting mislabeled examples, classifying minority-group samples, prioritizing points relevant for training and enabling scalable human auditing of relevant examples.
Holz, founder of AI art service Midjourney, on future images
Interview In 2008, David Holz co-founded a hardware peripheral firm called Leap Motion. He ran it until last year when he left to create Midjourey. Midjourney in its present form is a social network for creating AI-generated art from a text prompt โ type a word or phrase at the input prompt and you'll receive an interesting or perhaps wonderful image on screen after about a minute of computation. It's similar in some respects to OpenAI's DALL-E 2. Midjourney image of the sky and clouds, using the text prompt "All this useless beauty." Both are the result of large AI models trained on vast numbers of images. But Midjourney has its own distinctive style, as can be seen from this Twitter thread.
Gradient Health, Inc on LinkedIn: Data Requirements for FDA
Did you know: Representative Data We'd like to point out some key statistics from our last post on small study sizes. First of all, the question this article is trying to respond is about the prevalence and extent of small study effects in diagnostic imaging. Reach out to us to know how you can have quick access to millions of diverse medical imaging data and avoid data bias: https://lnkd.in/gVwPPXUB
World's first FLYING bike that can reach speeds of 62 mph and fly for 40 minutes to make US debut
A hoverbike that can travel at 62 miles per hour for up to 40 minutes made its U.S. debut this week at the North American Auto Show in Detroit. The flying bike is the work of Aerwins, a Delaware-based company that makes drones and unmanned vehicles. Although it conjures up futuristic Jetsons visions of s oaring high above New York City's notoriously clogged streets, you probably won't be riding the hoverbike out to John F. Kennedy Airport anytime soon. The Xturismo currently costs $777,000, although Aerwins says it will develop a smaller model next year, as well as an all-electric model in 2025 to sell for about $50,000. 'I feel like I'm literally 15-years-old and I just got out of Star Wars and I jumped on their bike,' Thad Scott, co-chair of the auto show, told Reuters.
Could artificial intelligence lead to genuine hiring bias?
There is a new bill to give the federal CXO councils more autonomy and structure. Senate lawmakers aim to give a boost to federal management with the Governmentwide Executive Councils Administration and Results Improvement Act. Sens. Gary Peters (D-Mich.) and Mike Braun (R-Ind.) said their legislation would make permanent and expand the Office of Executive Councils at the General Services Administration. The office helps run the federal chief information officer, chief financial officer and other similar councils. A key part of this effort would make the office independent from the normal duties and functions of GSA and more directly linked to the Office of Management and Budget, which leads the President's Management Agenda and other priorities.
What are People Talking about in #BlackLivesMatter and #StopAsianHate? Exploring and Categorizing Twitter Topics Emerging in Online Social Movements through the Latent Dirichlet Allocation Model
Tong, Xin, Li, Yixuan, Li, Jiayi, Bei, Rongqi, Zhang, Luyao
Minority groups have been using social media to organize social movements that create profound social impacts. Black Lives Matter (BLM) and Stop Asian Hate (SAH) are two successful social movements that have spread on Twitter that promote protests and activities against racism and increase the public's awareness of other social challenges that minority groups face. However, previous studies have mostly conducted qualitative analyses of tweets or interviews with users, which may not comprehensively and validly represent all tweets. Very few studies have explored the Twitter topics within BLM and SAH dialogs in a rigorous, quantified and data-centered approach. Therefore, in this research, we adopted a mixed-methods approach to comprehensively analyze BLM and SAH Twitter topics. We implemented (1) the latent Dirichlet allocation model to understand the top high-level words and topics and (2) open-coding analysis to identify specific themes across the tweets. We collected more than one million tweets with the #blacklivesmatter and #stopasianhate hashtags and compared their topics. Our findings revealed that the tweets discussed a variety of influential topics in depth, and social justice, social movements, and emotional sentiments were common topics in both movements, though with unique subtopics for each movement. Our study contributes to the topic analysis of social movements on social media platforms in particular and the literature on the interplay of AI, ethics, and society in general.