Goto

Collaborating Authors

 Government


Ceasing hate withMoH: Hate Speech Detection in Hindi-English Code-Switched Language

arXiv.org Artificial Intelligence

Social media has become a bedrock for people to voice their opinions worldwide. Due to the greater sense of freedom with the anonymity feature, it is possible to disregard social etiquette online and attack others without facing severe consequences, inevitably propagating hate speech. The current measures to sift the online content and offset the hatred spread do not go far enough. One factor contributing to this is the prevalence of regional languages in social media and the paucity of language flexible hate speech detectors. The proposed work focuses on analyzing hate speech in Hindi-English code-switched language. Our method explores transformation techniques to capture precise text representation. To contain the structure of data and yet use it with existing algorithms, we developed MoH or Map Only Hindi, which means "Love" in Hindi. MoH pipeline consists of language identification, Roman to Devanagari Hindi transliteration using a knowledge base of Roman Hindi words. Finally, it employs the fine-tuned Multilingual Bert and MuRIL language models. We conducted several quantitative experiment studies on three datasets and evaluated performance using Precision, Recall, and F1 metrics. The first experiment studies MoH mapped text's performance with classical machine learning models and shows an average increase of 13% in F1 scores. The second compares the proposed work's scores with those of the baseline models and offers a rise in performance by 6%. Finally, the third reaches the proposed MoH technique with various data simulations using the existing transliteration library. Here, MoH outperforms the rest by 15%. Our results demonstrate a significant improvement in the state-of-the-art scores on all three datasets.


Newsalyze: Effective Communication of Person-Targeting Biases in News Articles

arXiv.org Artificial Intelligence

Media bias and its extreme form, fake news, can decisively affect public opinion. Especially when reporting on policy issues, slanted news coverage may strongly influence societal decisions, e.g., in democratic elections. Our paper makes three contributions to address this issue. First, we present a system for bias identification, which combines state-of-the-art methods from natural language understanding. Second, we devise bias-sensitive visualizations to communicate bias in news articles to non-expert news consumers. Third, our main contribution is a large-scale user study that measures bias-awareness in a setting that approximates daily news consumption, e.g., we present respondents with a news overview and individual articles. We not only measure the visualizations' effect on respondents' bias-awareness, but we can also pinpoint the effects on individual components of the visualizations by employing a conjoint design. Our bias-sensitive overviews strongly and significantly increase bias-awareness in respondents. Our study further suggests that our content-driven identification method detects groups of similarly slanted news articles due to substantial biases present in individual news articles. In contrast, the reviewed prior work rather only facilitates the visibility of biases, e.g., by distinguishing left- and right-wing outlets.


How to Effectively Identify and Communicate Person-Targeting Media Bias in Daily News Consumption?

arXiv.org Artificial Intelligence

Slanted news coverage strongly affects public opinion. This is especially true for coverage on politics and related issues, where studies have shown that bias in the news may influence elections and other collective decisions. Due to its viable importance, news coverage has long been studied in the social sciences, resulting in comprehensive models to describe it and effective yet costly methods to analyze it, such as content analysis. We present an in-progress system for news recommendation that is the first to automate the manual procedure of content analysis to reveal person-targeting biases in news articles reporting on policy issues. In a large-scale user study, we find very promising results regarding this interdisciplinary research direction. Our recommender detects and reveals substantial frames that are actually present in individual news articles. In contrast, prior work rather only facilitates the visibility of biases, e.g., by distinguishing left- and right-wing outlets. Further, our study shows that recommending news articles that differently frame an event significantly improves respondents' awareness of bias.


Multilevel Stochastic Optimization for Imputation in Massive Medical Data Records

arXiv.org Machine Learning

Exploration and analysis of massive datasets has recently generated increasing interest in the research and development communities. It has long been a recognized problem that many datasets contain significant levels of missing numerical data. We introduce a mathematically principled stochastic optimization imputation method based on the theory of Kriging. This is shown to be a powerful method for imputation. However, its computational effort and potential numerical instabilities produce costly and/or unreliable predictions, potentially limiting its use on large scale datasets. In this paper, we apply a recently developed multi-level stochastic optimization approach to the problem of imputation in massive medical records. The approach is based on computational applied mathematics techniques and is highly accurate. In particular, for the Best Linear Unbiased Predictor (BLUP) this multi-level formulation is exact, and is also significantly faster and more numerically stable. This permits practical application of Kriging methods to data imputation problems for massive datasets. We test this approach on data from the National Inpatient Sample (NIS) data records, Healthcare Cost and Utilization Project (HCUP), Agency for Healthcare Research and Quality. Numerical results show the multi-level method significantly outperforms current approaches and is numerically robust. In particular, it has superior accuracy as compared with methods recommended in the recent report from HCUP on the important problem of missing data, which could lead to sub-optimal and poorly based funding policy decisions. In comparative benchmark tests it is shown that the multilevel stochastic method is significantly superior to recommended methods in the report, including Predictive Mean Matching (PMM) and Predicted Posterior Distribution (PPD), with up to 75% reductions in error.


COVID-19: Implications for business

#artificialintelligence

The Delta variant of the coronavirus spread to more countries in recent weeks, and the total number of cases officially logged soared past half a million per day. The global number of deaths is now about two-thirds as high as it was at the peak of the previous wave, in April of this year. As the virus spreads, the potential rises for a vaccine-resistant strain to emerge. Meanwhile, in poorer countries, vaccines are scarce, and most populations are little protected (exhibit).


US has lost AI race with China, ex-Pentagon chief says

#artificialintelligence

China has the competitive edge against the US in the field of artificial intelligence (AI), according to the Pentagon's former chief software officer. "We have no competing fighting chance against China in 15 to 20 years," Nicolas Chaillan said in an interview with London-based business newspaper, Financial Times. He called the current situation "a done deal," adding that, in his opinion, the race between China and the US was "already over." Chaillan predicted that China is heading for global dominance because of its advancements in the fields of artificial intelligence, machine learning and cyber capabilities, the Financial Times reported. He slammed US cyber defense capabilities as at "kindergarten level" in some government departments.


Clear the funding roadblock, AIIA urges on artificial intelligence

#artificialintelligence

More than $124 million in new funding for artificial intelligence research and industry development support allocated in the federal budget in May is still locked up inside the Industry department, with no clear signal on how and when it will be rolled out. The Australian Information Industry Association says Australia can't afford to sit on its hands in relation to the AI research and commercialisation โ€“ the industry is moving too fast, and the nation can't afford to fall behind. AIIA chief executive Ron Gauci says national capability in artificial intelligence is critical, because of the transformational impact that AI-based products and services are having across all industries. The AIIA has been pressing government for a funding allocation to drive commercialisation outcomes in the sector. The industry association had been told its "modest" proposal to bring together industry partners and state governments in a dollar-for-dollar funding arrangement with the Commonwealth had been agreed to.


UK schools will use facial recognition to speed up lunch payments

Engadget

Facial recognition may soon play a role in your child's lunch. The Financial Times reports that nine schools in the UK's North Ayrshire will start taking payments for canteen (aka cafeteria) lunches by scanning students' faces. The technology should help minimize touch during the pandemic, but is mainly meant to speed up transaction times. That could be important when you may have roughly 25 minutes to serve an entire school of hungry kids. Both the schools and system installer CRB Cunningham argued the systems would address privacy and security concerns.


AI and simulation tools to fight COVID-19

#artificialintelligence

In its on-going campaign to reveal the inner workings of the SARS-CoV-2 virus, the U.S. Department of Energy's (DOE) Argonne National Laboratory is leading efforts to couple artificial intelligence (AI) and cutting-edge simulation workflows to better understand biological observations and accelerate drug discovery. Argonne collaborated with academic and commercial research partners to achieve near real-time feedback between simulation and AI approaches to understand how two proteins in the SARS-CoV-2 viral genome, nsp10 and nsp16, interact to help the virus replicate and elude the host's immune system. The team achieved this milestone by coupling two distinct hardware platforms: Cerebras CS-1, a processor-packed silicon wafer deep learning accelerator; and ThetaGPU, an AI- and simulation-enabled extension of the Theta supercomputer, housed at the Argonne Leadership Computing Facility, a DOE Office of Science User Facility. To enable this capability, the team developed Stream-AI-MD, a novel application of the AI method called deep learning to drive adaptive molecular dynamics (MD) simulations in a streaming manner. Data from simulations is streamed from ThetaGPU onto the Cerebras CS-1 platform to simultaneously analyze how the two proteins interact.


Cornell Researchers Analyze Major Trends in Urban Tech

#artificialintelligence

A team of researchers at Cornell Tech, Cornell University's tech-focused research campus, has developed a forecast for how technologies like artificial intelligence could shape cities in the coming decade. After a year of work, the team released its first "Horizon Scan" report last week to discuss the potential risks and applications of recent advancements in urban tech. The forecast report predicts areas where the most radical and rapid changes in urban tech could take place, touching on topics such as "supercharged" smart city infrastructure, the use of sustainable building materials and machine learning in the public sector, among other areas of interest. The project was led by Anthony Townsend, urbanist in residence at the Jacobs Urban Tech Hub at Cornell Tech, who has spent years studying tech-related issues like the digital divide. He said the goal of the Horizon Scan was to create a road map "to make better decisions about applied research" in urban tech. Townsend said the need to weigh potential pros and cons of machine learning's applications in the public sector is a recurring factor in the report.