Africa
Detecting Languages Unintelligible to Multilingual Models through Local Structure Probes
Clouâtre, Louis, Parthasarathi, Prasanna, Zouaq, Amal, Chandar, Sarath
Providing better language tools for low-resource and endangered languages is imperative for equitable growth. Recent progress with massively multilingual pretrained models has proven surprisingly effective at performing zero-shot transfer to a wide variety of languages. However, this transfer is not universal, with many languages not currently understood by multilingual approaches. It is estimated that only 72 languages possess a "small set of labeled datasets" on which we could test a model's performance, the vast majority of languages not having the resources available to simply evaluate performances on. In this work, we attempt to clarify which languages do and do not currently benefit from such transfer. To that end, we develop a general approach that requires only unlabelled text to detect which languages are not well understood by a cross-lingual model. Our approach is derived from the hypothesis that if a model's understanding is insensitive to perturbations to text in a language, it is likely to have a limited understanding of that language. We construct a cross-lingual sentence similarity task to evaluate our approach empirically on 350, primarily low-resource, languages.
Cross-lingual Transfer Learning for Check-worthy Claim Identification over Twitter
Hasanain, Maram, Elsayed, Tamer
Misinformation spread over social media has become an undeniable infodemic. However, not all spreading claims are made equal. If propagated, some claims can be destructive, not only on the individual level, but to organizations and even countries. Detecting claims that should be prioritized for fact-checking is considered the first step to fight against spread of fake news. With training data limited to a handful of languages, developing supervised models to tackle the problem over lower-resource languages is currently infeasible. Therefore, our work aims to investigate whether we can use existing datasets to train models for predicting worthiness of verification of claims in tweets in other languages. We present a systematic comparative study of six approaches for cross-lingual check-worthiness estimation across pairs of five diverse languages with the help of Multilingual BERT (mBERT) model. We run our experiments using a state-of-the-art multilingual Twitter dataset. Our results show that for some language pairs, zero-shot cross-lingual transfer is possible and can perform as good as monolingual models that are trained on the target language. We also show that in some languages, this approach outperforms (or at least is comparable to) state-of-the-art models.
Hibikino-Musashi@Home 2018 Team Description Paper
Ishida, Yutaro, Hori, Sansei, Tanaka, Yuichiro, Yoshimoto, Yuma, Hashimoto, Kouhei, Iwamoto, Gouki, Aratani, Yoshiya, Yamashita, Kenya, Ishimoto, Shinya, Hitaka, Kyosuke, Yamaguchi, Fumiaki, Miyoshi, Ryuhei, Honda, Kentaro, Abe, Yushi, Kato, Yoshitaka, Morie, Takashi, Tamukoh, Hakaru
Our team, Hibikino-Musashi@Home (the shortened name is HMA), was founded in 2010. It is based in the Kitakyushu Science and Research Park, Japan. We have participated in the RoboCup@Home Japan open competition open platform league every year since 2010. Moreover, we participated in the RoboCup 2017 Nagoya as open platform league and domestic standard platform league teams. Currently, the Hibikino-Musashi@Home team has 20 members from seven different laboratories based in the Kyushu Institute of Technology. In this paper, we introduce the activities of our team and the technologies.
Graph representation learning for street networks
Streets networks provide an invaluable source of information about the different temporal and spatial patterns emerging in our cities. These streets are often represented as graphs where intersections are modelled as nodes and streets as links between them. Previous work has shown that raster representations of the original data can be created through a learning algorithm on low-dimensional representations of the street networks. In contrast, models that capture high-level urban network metrics can be trained through convolutional neural networks. However, the detailed topological data is lost through the rasterisation of the street network. The models cannot recover this information from the image alone, failing to capture complex street network features. This paper proposes a model capable of inferring good representations directly from the street network. Specifically, we use a variational autoencoder with graph convolutional layers and a decoder that outputs a probabilistic fully-connected graph to learn latent representations that encode both local network structure and the spatial distribution of nodes. We train the model on thousands of street network segments and use the learnt representations to generate synthetic street configurations. Finally, we proposed a possible application to classify the urban morphology of different network segments by investigating their common characteristics in the learnt space.
Google wants AI language in 1,000 dialects
Google last week unpacked a host of critical new artificial intelligence (AI) advancements, particularly in the areas around generative AI such as language analysis, text-to-video, and assisted writing capabilities. Along with other enticing developments around ethical AI, Google revealed it would be working on AI in at least one thousand languages going forward. At the Google AI event held at the company's Pier 57 offices in New York City on November 2, the global technology titan highlighted areas of focus to include democratizing AI development pathways, "building for everyone" responsible, controlled models which can identify generative AI -- unsupervised artificial intelligence learning algorithms that can create new digital audio, text, imagery, video, or code -- along with advancing language translation, disaster management, and health AI. Google shared the first rendering of a video that shares both of the company's complementary text-to-video research approaches -- Imagen Video and Phenaki. The interest of global communities was piqued by the announcement that language abilities are being developed using AI for the world's one thousand most spoken languages, showcasing another major area of artificial intelligence focus as tech giants compete to dominate the internet's next battleground.
Common names in Burkina Faso, West-Africa
Burkina Faso is a multi-cultural and diverse country with a rich history. In this article, we explore how personal names can be interpreted to reflect regional, ethnic appartenance within the country. Then we illustrate how the use of a personal name can affect a black-box Artificial Intelligence – such as OpenAI's DALL-E. This is a first article in our series of blog posts with tag #thisnamedpersondoesnotexist.
Blockchain and AI Technology: Benefiting the Ordinary Citizen Part 1
Blockchain and AI, particularly machine learning, are two quite recent revolutionary technologies that are being adopted by Governments and Businesses in all sorts of ways. In a series of articles (in 4 parts) I reflect in what ways these technologies have the potential to impact the lives of ordinary citizens. A lot of discussion has been going on about how blockchain and AI affects governments and businesses. If AI has been around for quite a long time, the more recent Blockchain technology, has taken the world by storm, as being a database system that provides us with a simple protocol that allows transactions to be simultaneously anonymous and secure, peer-to-peer, immediate and in constant flow. The beautiful promise of blockchain is that is distributes the trust which is currently allocated to centralized, large and powerful intermediaries, to a large global network of people engaged in massive collaboration, facilitated by clever coding and cryptography.
JARVIS Invest rolls out its AI in investing campaign
Artificial intelligence (AI)-based investment advisory platform JARVIS Invest has launched its first brand campaign'AI in investing'. Through the campaign, the brand aims to highlight the benefits of AI and its application in stock advisory services. As per the company, the campaign will run across various digital platforms for five months using a programmatic approach. JARVIS Invest has recently crossed assets under advisory of over Rs 100 crores and the AI tool'JARVIS' has also been able to successfully predict the recent market crash of March 2020 and January 2022, Sumit Chanda, CEO and founder, Jarvis Invest said. "While investors have been receptive to the idea of AI, there is still certain resistance from the prospects of trusting AI completely and using AI in their journey of wealth creation. This triggered us to devise a campaign that could address this apprehension. Through this campaign, early-age investors will now be able to gain confidence about creating their custom portfolios using AI and fulfil their financial goals," he added.
These A.I.-Generated Images Hang in a Gallery--but Are They Art?
When it comes to creativity, is artificial intelligence a powerful new tool or an existential threat? A San Francisco gallery is taking on this question in a new exhibition: "Artificial Imagination" features eight artists who used A.I. image generators to create the pieces on display. The artists' methods vary: Some fed their A.I. tool of choice phrases to generate their entire piece, while others created illustrations or sculptures based on the tool's recommendations. The show is on view at bitforms' West Coast gallery through the end of the year. From robots that make their own art to image-generation tools that mimick history's greatest painters, A.I. is quickly permeating creative spaces--and generating lots of questions.
Maximum likelihood recursive state estimation in state-space models: A new approach based on statistical analysis of incomplete data
This paper revisits the work of Rauch et al. (1965) and develops a novel method for recursive maximum likelihood particle filtering for general state-space models. The new method is based on statistical analysis of incomplete observations of the systems. Score function and conditional observed information of the incomplete observations/data are introduced and their distributional properties are discussed. Some identities concerning the score function and information matrices of the incomplete data are derived. Maximum likelihood estimation of state-vector is presented in terms of the score function and observed information matrices. In particular, to deal with nonlinear state-space, a sequential Monte Carlo method is developed. It is given recursively by an EM-gradient-particle filtering which extends the work of Lange (1995) for state estimation. To derive covariance matrix of state-estimation errors, an explicit form of observed information matrix is proposed. It extends Louis (1982) general formula for the same matrix to state-vector estimation. Under (Neumann) boundary conditions of state transition probability distribution, the inverse of this matrix coincides with the Cramer-Rao lower bound on the covariance matrix of estimation errors of unbiased state-estimator. In the case of linear models, the method shows that the Kalman filter is a fully efficient state estimator whose covariance matrix of estimation error coincides with the Cramer-Rao lower bound. Some numerical examples are discussed to exemplify the main results.