Africa
Few-shot Learning with Multilingual Language Models
Lin, Xi Victoria, Mihaylov, Todor, Artetxe, Mikel, Wang, Tianlu, Chen, Shuohui, Simig, Daniel, Ott, Myle, Goyal, Naman, Bhosale, Shruti, Du, Jingfei, Pasunuru, Ramakanth, Shleifer, Sam, Koura, Punit Singh, Chaudhary, Vishrav, O'Horo, Brian, Wang, Jeff, Zettlemoyer, Luke, Kozareva, Zornitsa, Diab, Mona, Stoyanov, Veselin, Li, Xian
Large-scale autoregressive language models such as GPT-3 are few-shot learners that can perform a wide range of language tasks without fine-tuning. While these models are known to be able to jointly represent many different languages, their training data is dominated by English, potentially limiting their cross-lingual generalization. In this work, we train multilingual autoregressive language models on a balanced corpus covering a diverse set of languages, and study their few- and zero-shot learning capabilities in a wide range of tasks. Our largest model with 7.5 billion parameters sets new state of the art in few-shot learning in more than 20 representative languages, outperforming GPT-3 of comparable size in multilingual commonsense reasoning (with +7.4% absolute accuracy improvement in 0-shot settings and +9.4% in 4-shot settings) and natural language inference (+5.4% in each of 0-shot and 4-shot settings). On the FLORES-101 machine translation benchmark, our model outperforms GPT-3 on 171 out of 182 translation directions with 32 training examples, while surpassing the official supervised baseline in 45 directions. We present a detailed analysis of where the model succeeds and fails, showing in particular that it enables cross-lingual in-context learning on some tasks, while there is still room for improvement on surface form robustness and adaptation to tasks that do not have a natural cloze form. Finally, we evaluate our models in social value tasks such as hate speech detection in five languages and find it has limitations similar to comparable sized GPT-3 models.
Autonomous Weapons Are Here, but the World Isn't Ready for Them
This may be remembered as the year when the world learned that lethal autonomous weapons had moved from a futuristic worry to a battlefield reality. It's also the year when policymakers failed to agree on what to do about it. On Friday, 120 countries participating in the United Nations' Convention on Certain Conventional Weapons could not agree on whether to limit the development or use of lethal autonomous weapons. Instead, they pledged to continue and "intensify" discussions. "It's very disappointing, and a real missed opportunity," says Neil Davison, senior scientific and policy adviser at the International Committee of the Red Cross, a humanitarian organization based in Geneva.
Hidden Pentagon records reveal patterns of failure in deadly U.S. airstrikes
Shortly before 3 a.m. on July 19, 2016, U.S. Special Operations forces bombed what they believed were three Islamic State (IS) group "staging areas" on the outskirts of Tokhar, a riverside hamlet in northern Syria. They reported 85 fighters killed. In fact, they hit houses far from the front line, where farmers, their families and other local people sought nighttime sanctuary from bombing and gunfire. More than 120 villagers were killed. In early 2017 in Iraq, an American war plane struck a dark-colored vehicle, believed to be a car bomb, stopped at an intersection in the Wadi Hajar neighborhood of West Mosul. Actually, the car had been bearing not a bomb but a man named Majid Mahmoud Ahmed, his wife and their two children, who were fleeing the fighting nearby. They and three other civilians were killed. In November 2015, after observing a man dragging an "unknown heavy object" into an IS "defensive fighting position," U.S. forces struck a building in Ramadi, Iraq. A military review found that the object was actually "a person of small stature" -- a child -- who died in the strike. None of these deadly failures resulted in a finding of wrongdoing. These cases are drawn from a hidden Pentagon archive of the American air war in the Middle East since 2014. The trove of documents -- the military's own confidential assessments of more than 1,300 reports of civilian casualties, obtained by The New York Times -- lays bare how the air war has been marked by deeply flawed intelligence, rushed and often imprecise targeting and the deaths of thousands of civilians, many of them children, a sharp contrast to the U.S. government's image of war waged by all-seeing drones and precision bombs. The documents show, too, that despite the Pentagon's highly codified system for examining civilian casualties, pledges of transparency and accountability have given way to opacity and impunity. In only a handful of cases were the assessments made public. Not a single record provided includes a finding of wrongdoing or disciplinary action. Fewer than a dozen condolence payments were made, even though many survivors were left with disabilities requiring expensive medical care. Documented efforts to identify root causes or lessons learned are rare. The air campaign represents a fundamental transformation of warfare that took shape in the final years of the Obama administration, amid the deepening unpopularity of the forever wars that had claimed more than 6,000 American service members. The United States traded many of its boots on the ground for an arsenal of aircraft directed by controllers sitting at computers, often thousands of kilometers away. President Barack Obama called it "the most precise air campaign in history." This was the promise: America's "extraordinary technology" would allow the military to kill the right people while taking the greatest possible care not to harm the wrong ones. The IS caliphate ultimately crumbled under the weight of American bombing.
Offline Pre-trained Multi-Agent Decision Transformer: One Big Sequence Model Tackles All SMAC Tasks
Meng, Linghui, Wen, Muning, Yang, Yaodong, Le, Chenyang, Li, Xiyun, Zhang, Weinan, Wen, Ying, Zhang, Haifeng, Wang, Jun, Xu, Bo
Offline reinforcement learning leverages previously-collected offline datasets to learn optimal policies with no necessity to access the real environment. Such a paradigm is also desirable for multi-agent reinforcement learning (MARL) tasks, given the increased interactions among agents and with the enviroment. Yet, in MARL, the paradigm of offline pre-training with online fine-tuning has not been studied, nor datasets or benchmarks for offline MARL research are available. In this paper, we facilitate the research by providing large-scale datasets, and use them to examine the usage of the Decision Transformer in the context of MARL. We investigate the generalisation of MARL offline pre-training in the following three aspects: 1) between single agents and multiple agents, 2) from offline pretraining to the online fine-tuning, and 3) to that of multiple downstream tasks with few-shot and zero-shot capabilities. We start by introducing the first offline MARL dataset with diverse quality levels based on the StarCraftII environment, and then propose the novel architecture of multi-agent decision transformer (MADT) for effective offline learning. MADT leverages transformer's modelling ability of sequence modelling and integrates it seamlessly with both offline and online MARL tasks. A crucial benefit of MADT is that it learns generalizable policies that can transfer between different types of agents under different task scenarios. On StarCraft II offline dataset, MADT outperforms the state-of-the-art offline RL baselines. When applied to online tasks, the pre-trained MADT significantly improves sample efficiency, and enjoys strong performance both few-short and zero-shot cases. To our best knowledge, this is the first work that studies and demonstrates the effectiveness of offline pre-trained models in terms of sample efficiency and generalisability enhancements in MARL.
3 big problems with datasets in AI and machine learning
Datasets fuel AI models like gasoline (or electricity, as the case may be) fuels cars. Whether they're tasked with generating text, recognizing objects, or predicting a company's stock price, AI systems "learn" by sifting through countless examples to discern patterns in the data. For example, a computer vision system can be trained to recognize certain types of apparel, like coats and scarfs, by looking at different images of that clothing. Beyond developing models, datasets are used to test trained AI systems to ensure they remain stable -- and measure overall progress in the field. Models that top the leaderboards on certain open source benchmarks are considered state of the art (SOTA) for that particular task.
UN talks fail to open negotiations on 'killer robots'
Country officials and campaigners have expressed disappointment after United Nations talks on autonomous weapons systems – known as "killer robots" – stopped short of launching negotiations into an international treaty to govern their use following opposition from manufacturing states. Unlike existing semi-autonomous weapons such as drones, fully-autonomous weapons have no human-operated "kill switch" and instead leave decisions over life and death to sensors, software and machine processes. The regulation of the industry has taken on new urgency since a UN panel report in March said the first autonomous drone attack may have occurred in Libya. This week, UN Secretary-General Antonio Guterres encouraged the 125 parties to the Convention on Certain Conventional Weapons (CCW) to come up with an "ambitious plan" on new rules. But on Friday, the Sixth Review Conference of the CCW failed to schedule further talks around the development and use of the Lethal Autonomous Weapon Systems, or LAWS.
Mobile Artificial Intelligence (AI) Market to Generate Massive USD 29.34 billion by 2027 - Digital Journal
"The Global Mobile Artificial Intelligence (AI) Market analysis provides a high-level summary of classification, competition, and strategic actions taken in recent years. For a global scenario, the global Mobile Artificial Intelligence (AI) market report provides historical details, future forecasts, and market size. The Mobile Artificial Intelligence (AI) report displays important product developments and tracks recent acquisitions, mergers and research in this industry by the key players. Mobile Artificial Intelligence (AI) report also puts light on the company market share analysis and key company profiles which are the major aspects of competitive analysis. Being a verified and reliable source of information, this market research report offers a telescopic view of the existing market trends, emerging products, situations and opportunities that drives the business in the right direction of success.
AI 50 2021: America's Most Promising Artificial Intelligence Companies
The Covid-19 pandemic was devastating for many industries, but it only accelerated the use of artificial intelligence across the U.S. economy. Amid the crisis, companies scrambled to create new services for remote workers and students, beef up online shopping and dining options, make customer call centers more efficient and speed development of important new drugs. Even as applications of machine learning and perception platforms become commonplace, a thick layer of hype and fuzzy jargon clings to AI-enabled software.That makes it tough to identify the most compelling companies in the space--especially those finding new ways to use AI that create value by making humans more efficient, not redundant. With this in mind, Forbes has partnered with venture firms Sequoia Capital and Meritech Capital to create our third annual AI 50, a list of private, promising North American companies that are using artificial intelligence in ways that are fundamental to their operations. To be considered, businesses must be privately-held and utilizing machine learning (where systems learn from data to improve on tasks), natural language processing (which enables programs to "understand" written or spoken language) or computer vision (which relates to how machines "see"). AI companies incubated at, largely funded through or acquired by large tech, manufacturing or industrial firms aren't eligible for consideration. Our list was compiled through a submission process open to any AI company in the U.S. and Canada. The application asked companies to provide details on their technology, business model, customers and financials like funding, valuation and revenue history (companies had the option to submit information confidentially, to encourage greater transparency). Forbes received several hundred entries, of which nearly 400 qualified for consideration. From there, our data partners applied an algorithm to identify 100 companies with the highest quantitative scores--and that also made diversity a priority. Next, a panel of expert AI judges evaluated the finalists to find the 50 most compelling companies (they were precluded from judging companies in which they have a vested interest). Among trends this year are what Sequoia Capital's Konstantine Buhler calls AI workbench companies--building of platforms tailored to different enterprises, including Dataiku, DataRobot Domino Data and Databricks.
Context-self contrastive pretraining for crop type semantic segmentation
Tarasiou, Michail, Guler, Riza Alp, Zafeiriou, Stefanos
In this paper, we propose a fully supervised pre-training scheme based on contrastive learning particularly tailored to dense classification tasks. The proposed Context-Self Contrastive Loss (CSCL) learns an embedding space that makes semantic boundaries pop-up by use of a similarity metric between every location in a training sample and its local context. For crop type semantic segmentation from Satellite Image Time Series (SITS) we find performance at parcel boundaries to be a critical bottleneck and explain how CSCL tackles the underlying cause of that problem, improving the state-of-the-art performance in this task. Additionally, using images from the Sentinel-2 (S2) satellite missions we compile the largest, to our knowledge, SITS dataset densely annotated by crop type and parcel identities, which we make publicly available together with the data generation pipeline. Using that data we find CSCL, even with minimal pre-training, to improve all respective baselines and present a process for semantic segmentation at super-resolution for obtaining crop classes at a more granular level. The code and instructions to download the data can be found in https://github.com/michaeltrs/DeepSatModels.