Government
Deciphering Hate: Identifying Hateful Memes and Their Targets
Hossain, Eftekhar, Sharif, Omar, Hoque, Mohammed Moshiul, Preum, Sarah M.
Internet memes have become a powerful means for individuals to express emotions, thoughts, and perspectives on social media. While often considered as a source of humor and entertainment, memes can also disseminate hateful content targeting individuals or communities. Most existing research focuses on the negative aspects of memes in high-resource languages, overlooking the distinctive challenges associated with low-resource languages like Bengali (also known as Bangla). Furthermore, while previous work on Bengali memes has focused on detecting hateful memes, there has been no work on detecting their targeted entities. To bridge this gap and facilitate research in this arena, we introduce a novel multimodal dataset for Bengali, BHM (Bengali Hateful Memes). The dataset consists of 7,148 memes with Bengali as well as code-mixed captions, tailored for two tasks: (i) detecting hateful memes, and (ii) detecting the social entities they target (i.e., Individual, Organization, Community, and Society). To solve these tasks, we propose DORA (Dual cO attention fRAmework), a multimodal deep neural network that systematically extracts the significant modality features from the memes and jointly evaluates them with the modality-specific features to understand the context better. Our experiments show that DORA is generalizable on other low-resource hateful meme datasets and outperforms several state-of-the-art rivaling baselines.
From Pixels to Predictions: Spectrogram and Vision Transformer for Better Time Series Forecasting
Zeng, Zhen, Kaur, Rachneet, Siddagangappa, Suchetha, Balch, Tucker, Veloso, Manuela
Time series forecasting plays a crucial role in decision-making across various domains, but it presents significant challenges. Recent studies have explored image-driven approaches using computer vision models to address these challenges, often employing lineplots as the visual representation of time series data. In this paper, we propose a novel approach that uses time-frequency spectrograms as the visual representation of time series data. We introduce the use of a vision transformer for multimodal learning, showcasing the advantages of our approach across diverse datasets from different domains. To evaluate its effectiveness, we compare our method against statistical baselines (EMA and ARIMA), a state-of-the-art deep learning-based approach (DeepAR), other visual representations of time series data (lineplot images), and an ablation study on using only the time series as input. Our experiments demonstrate the benefits of utilizing spectrograms as a visual representation for time series data, along with the advantages of employing a vision transformer for simultaneous learning in both the time and frequency domains.
Advancing multivariate time series similarity assessment: an integrated computational approach
Tonle, Franck, Tonnang, Henri, Ndadji, Milliam, Tchendji, Maurice, Nzeukou, Armand, Senagi, Kennedy, Niassy, Saliou
Data mining, particularly the analysis of multivariate time series data, plays a crucial role in extracting insights from complex systems and supporting informed decision-making across diverse domains. However, assessing the similarity of multivariate time series data presents several challenges, including dealing with large datasets, addressing temporal misalignments, and the need for efficient and comprehensive analytical frameworks. To address all these challenges, we propose a novel integrated computational approach known as Multivariate Time series Alignment and Similarity Assessment (MTASA). MTASA is built upon a hybrid methodology designed to optimize time series alignment, complemented by a multiprocessing engine that enhances the utilization of computational resources. This integrated approach comprises four key components, each addressing essential aspects of time series similarity assessment, thereby offering a comprehensive framework for analysis. MTASA is implemented as an open-source Python library with a user-friendly interface, making it accessible to researchers and practitioners. To evaluate the effectiveness of MTASA, we conducted an empirical study focused on assessing agroecosystem similarity using real-world environmental data. The results from this study highlight MTASA's superiority, achieving approximately 1.5 times greater accuracy and twice the speed compared to existing state-of-the-art integrated frameworks for multivariate time series similarity assessment. It is hoped that MTASA will significantly enhance the efficiency and accessibility of multivariate time series analysis, benefitting researchers and practitioners across various domains. Its capabilities in handling large datasets, addressing temporal misalignments, and delivering accurate results make MTASA a valuable tool for deriving insights and aiding decision-making processes in complex systems.
Regulating Chatbot Output via Inter-Informational Competition
The advent of ChatGPT has sparked over a year of regulatory frenzy. However, few existing studies have rigorously questioned the assumption that, if left unregulated, AI chatbot's output would inflict tangible, severe real harm on human affairs. Most researchers have overlooked the critical possibility that the information market itself can effectively mitigate these risks and, as a result, they tend to use regulatory tools to address the issue directly. This Article develops a yardstick for reevaluating both AI-related content risks and corresponding regulatory proposals by focusing on inter-informational competition among various outlets. The decades-long history of regulating information and communications technologies indicates that regulators tend to err too much on the side of caution and to put forward excessive regulatory measures when encountering the uncertainties brought about by new technologies. In fact, a trove of empirical evidence has demonstrated that market competition among information outlets can effectively mitigate most risks and that overreliance on regulation is not only unnecessary but detrimental, as well. This Article argues that sufficient competition among chatbots and other information outlets in the information marketplace can sufficiently mitigate and even resolve most content risks posed by generative AI technologies. This renders certain loudly advocated regulatory strategies, like mandatory prohibitions, licensure, curation of datasets, and notice-and-response regimes, truly unnecessary and even toxic to desirable competition and innovation throughout the AI industry. Ultimately, the ideas that I advance in this Article should pour some much-needed cold water on the regulatory frenzy over generative AI and steer the issue back to a rational track.
Reddit's Sale of User Data for AI Training Draws FTC Inquiry
Reddit said ahead of its IPO next week that licensing user posts to Google and others for AI projects could bring in 203 million of revenue over the next few years. The community-driven platform was forced to disclose Friday that US regulators already have questions about that new line of business. In a regulatory filing, Reddit said that it received a letter from the US Federal Trade Commision on Thursday asking about "our sale, licensing, or sharing of user-generated content with third parties to train AI models." The FTC, the US government's primary antitrust regulator, has the power to sanction companies found to engage in unfair or deceptive trade practices. Reddit isn't alone in trying to make a buck off licensing data, including that generated by users, for AI.
Governments Setting Limits on AI
The Biden Administration's actions came on the heels of the European Union, which last June passed the landmark Artificial Intelligence Act, moving a step closer to formally adopting the first-of-its-kind set of comprehensive rules around regulating AI. The AI Act, which was expected to be adopted early this year, sets four classifications for AI risk, ranging from minimal to unacceptable. Technology classified as an unacceptable risk, for example, would include systems that judge people based on a behavior known as social scoring, along with predictive policing tools, and would be banned. There also will be an EU AI board to oversee the implementation and uniform application of the regulations, which will build on existing GDPR and Intellectual Property legislation. The AI Act "is the first comprehensive regulation addressing the risks of artificial intelligence through a set of obligations and requirements that intend to safeguard the health, safety and fundamental rights of EU citizens and beyond, and is expected to have an outsized impact on AI governance worldwide," wrote Mia Hoffmann, a research fellow at the Center for Security and Emerging Technology (CSET) at Georgetown University.
The Problem With em Dune: Part Two /em
I have questions about Denis Villeneuve's Dune: Part Two. If the Fremen have lasers, why don't they just shoot the sand harvesters and run away? Why don't they use their sandworms until the last battle? Wouldn't it make more sense to fight the other great houses on Arrakis itself, where they have sandworms, rather than board ships off-world to go off to war? If Paul (Timothรฉe Chalamet) has to invade the galaxy at the end, why bother marrying the daughter of the emperor he just deposed?
Learning a Distance Metric from a Network
Many real-world networks are described by both connectivity information and features for every node. To better model and understand these networks, we present structure preserving metric learning (SPML), an algorithm for learning a Mahalanobis distance metric from a network such that the learned distances are tied to the inherent connectivity structure of the network. Like the graph embedding algorithm structure preserving embedding, SPML learns a metric which is structure preserving, meaning a connectivity algorithm such as k-nearest neighbors will yield the correct connectivity when applied using the distances from the learned metric. We show a variety of synthetic and real-world experiments where SPML predicts link patterns from node features more accurately than standard techniques. We further demonstrate a method for optimizing SPML based on stochastic gradient descent which removes the running-time dependency on the size of the network and allows the method to easily scale to networks of thousands of nodes and millions of edges.
Africa's push to regulate AI starts now
Now, the African Union--made up of 55 member nations--is preparing an ambitious AI policy that envisions an Africa-centric path for the development and regulation of this emerging technology. But debates on when AI regulation is warranted and concerns about stifling innovation could pose a roadblock, while a lack of AI infrastructure could hold back the technology's adoption. "We're seeing a growth of AI in the continent; it's really important there be set rules in place to govern these technologies," says Chinasa T. Okolo, a fellow in the Center for Technology Innovation at Brookings, whose research focuses on AI governance and policy development in Africa. Some African countries have already begun to formulate their own legal and policy frameworks for AI. Seven have developed national AI policies and strategies, which are currently at different stages of implementation.
Giant volcano 'hidden in plain sight' discovered on Mars, scientists say
Scientists say they have discovered a giant volcano hidden in plain sight on Mars. The volcano, temporarily named the Noctis, spans 280 miles wide and was discovered alongside a buried ice glacier to the east of Mars, near the red-planet's equator, scientists revealed at the 55th Lunar and Planetary Science Conference held in Texas on Wednesday. Scientists said the 29,600-foot-high volcano was active from ancient through recent times and with possible remnants of glacier ice near its base. They say its discovery points to an exciting new place to search for life and a potential destination for future robotic and human exploration. The findings were detailed in a new study by the SETI Institute and the Mars Institute based at NASA Ames Research Centre. Scientists have discovered a gigantic volcano on Mars that spans 280 miles wide and and nearly 30,000 feet high near the red-planet's equator, (NASA/USGS Mars globe.