Government
niderhoff/nlp-datasets
Most stuff here is just raw unstructured text data, if you are looking for annotated corpora or Treebanks refer to the sources at the bottom. Blog Authorship Corpus: consists of the collected posts of 19,320 bloggers gathered from blogger.com in August 2004. Amazon Fine Food Reviews [Kaggle]: consists of 568,454 food reviews Amazon users left up to October 2012. ASAP Automated Essay Scoring [Kaggle]: For this competition, there are eight essay sets. Each of the sets of essays was generated from a single prompt.
The Global AI Index - Tortoise
Artificial intelligence is poised to transform the way we work, learn and live. Across the globe, businesses, governments and the public at large are already having to adapt to the rapid development of these technologies. The Global AI Index analyses how 54 countries are driving and adapting to AI's accelerating development through three pillars; investment, innovation and implementation. Here is the Index in full. Use the toggle to switch between our Index's ranks โ where countries stand โ or score โ how far or close they are to each other.
Prestigious Pyongyang university teaching specialist Japanese language, literature courses
Kim Il Sung University in the spring of 2017 set up specialist Japanese language and literature courses, it was learned Saturday from the university. The training course for Japanese researchers was established at the prestigious institution in the capital, Pyongyang, at a time when North Korea was repeatedly testing nuclear weapons and launching ballistic missiles, which continued until the fall of 2017 and led to heightened tensions with the United States. There is a possibility that it was judged necessary to strengthen the development of such experts in view of future diplomacy with Japan. Japan and North Korea have no diplomatic relations. The Department of Japanese Language and Literature was established in the university's Faculty of Foreign Languages and Literature.
3 Ways to Discover AI Trends in Any Sector Emerj
Daniel Faggella is the founder and CEO at Emerj. Called upon by the United Nations, World Bank, INTERPOL, and many global enterprises, Daniel is a sought-after expert on the competitive strategy implications of AI for business and government leaders. Business leaders, managers, and consultants with an eye on AI aren't just trying to learn what AI can do, they're trying to discover ways to gain an AI advantage. For this reason, discovering AI trends can be particularly important. Most of the work that we do with our AI Capability Map services is about finding trends in quantitative data โ which requires hundreds of hours of expert research, and established frameworks for interpreting and categorizing data for insight.
CIMON, the AI-powered robot, launches a new era in space travel
Outer space is unfriendly to humans. We can only exist in a contained environment with its own air supply, so that means working in very close quarters. Zero gravity means things don't stay where you put them. Despite these constraints, astronauts on the International Space Station (ISS) must be highly productive on a very tight time schedule. They conduct hundreds of experiments they've memorized and trained for on the ground, but now in zero-gravity.
New AI-driven technology for breast cancer screening
This included the technology ProFound AI for Digital Breast Tomosynthesis (DBT), which is said to be the first artificial intelligence software for DBT to be approved by the U.S. Food and Drug Administration (FDA). Also on offer at the event were medical software solutions designed for 2D mammography and to assess breast density. During the meeting, the iCAD unveiled its vision for future technologies. This predictive aspect included technologies that should enable clinicians to more easily interpret patients' earlier images and prospective breast cancer risk assessment to form a clearer picture of the specific patient's condition. Clinical data from a large reader study involving ProFound AI for DBT were recently published in the journal Radiology: Artificial Intelligence ("Improving Accuracy and Efficiency with Concurrent Use of Artificial Intelligence for Digital Breast Tomosynthesis").
141 Cybersecurity Predictions For 2020
Serial cybersecurity entrepreneur Shlomo Kramer said in a 2005 interview that cybersecurity is "a bit like Alice in Wonderland" where you run as fast as you can only to stay in place. In 2020, to paraphrase the second part of the Red Queen's observation (actually from Through the Looking Glass), if you wish to stay ahead of cyber criminals, you must run twice--or ten times--as fast as that. The 141 predictions listed here reveal the state-of-mind of key participants in the cybersecurity defense industry and highlight all that's hot today. The future is murky, but we know for sure that on January 1, 2020, the California Consumer Privacy Act (CCPA) will go into effect; that the U.S. presidential election will take place on November 3, 2020; and that on October 1, 2020, if you "wish to fly on commercial aircrafts or access federal facilities" in the U.S., you must have a REAL ID compliant card. Other than these known events, the crystal balls of the participants in this survey warn us ...
Group Fairness in Bandit Arm Selection
Schumann, Candice, Lang, Zhi, Mattei, Nicholas, Dickerson, John P.
We consider group fairness in the contextual bandit setting. Here, a sequential decision maker must choose at each time step an arm to pull from a finite set of arms, after observing some context for each of the potential arm pulls. Additionally, arms are partitioned into m sensitive groups based on some protected feature (e.g., age, race, or socio-economic status). Despite the fact that there may be differences in expected payout between the groups, we may wish to ensure some form of fairness between picking arms from the various groups. In this work, we explore two definitions of fairness: equal group probability, wherein the probability of pulling an arm from any of the protected groups is the same; and proportional parity, wherein the probability of choosing an arm from a particular group is proportional to the size of that group. We provide a novel algorithm that can accommodate these notions of fairness and provide bounds on the regret for our algorithm. We test our algorithms on a hypothetical intervention setting wherein we want to allocate resources across protected groups.
Hybrid Compositional Reasoning for Reactive Synthesis from Finite-Horizon Specifications
Bansal, Suguman, Li, Yong, Tabajara, Lucas M., Vardi, Moshe Y.
LTLf synthesis is the automated construction of a reactive system from a high-level description, expressed in LTLf, of its finite-horizon behavior. So far, the conversion of LTLf formulas to deterministic finite-state automata (DFAs) has been identified as the primary bottleneck to the scalabity of synthesis. Recent investigations have also shown that the size of the DFA state space plays a critical role in synthesis as well. Therefore, effective resolution of the bottleneck for synthesis requires the conversion to be time and memory performant, and prevent state-space explosion. Current conversion approaches, however, which are based either on explicit-state representation or symbolic-state representation, fail to address these necessities adequately at scale: Explicit-state approaches generate minimal DFA but are slow due to expensive DFA minimization. Symbolic-state representations can be succinct, but due to the lack of DFA minimization they generate such large state spaces that even their symbolic representations cannot compensate for the blow-up. This work proposes a hybrid representation approach for the conversion. Our approach utilizes both explicit and symbolic representations of the state-space, and effectively leverages their complementary strengths. In doing so, we offer an LTLf to DFA conversion technique that addresses all three necessities, hence resolving the bottleneck. A comprehensive empirical evaluation on conversion and synthesis benchmarks supports the merits of our hybrid approach.
VAT tax gap prediction: a 2-steps Gradient Boosting approach
Tagliaferri, Giovanna, Scacciatelli, Daria, Di Loro, Pierfrancesco Alaimo
Tax evasion is the illegal non-payment of taxes by individuals, corporations, and trusts. It results in a loss of state revenue that can undermine the effectiveness of government policies. One measure of tax evasion is the so-called tax gap: the difference between the income that should be reported to the tax authorities and the amount actually reported. However, economists lack a robust method for estimating the tax gap through a bottom-up approach based on fiscal audits. This is difficult because the declared tax base is available on the whole population but the income reported to the tax authorities is generally available only on a small, non-random sample of audited units. This induces a selection bias which invalidates standard statistical methods. Here, we use machine learning based on a 2-steps Gradient Boosting model, to correct for the selection bias without requiring any strong assumption on the distribution. We use our method to estimate the Italian VAT Gap related to individual firms based on information gathered from administrative sources. Our algorithm estimates the potential VAT turnover of Italian individual firms for the fiscal year 2011 and suggests that the tax gap is about 30% of the total potential tax base. Comparisons with other methods show our technique offers a significant improvement in predictive performance.