Goto

Collaborating Authors

 Information Extraction


Identifying Offensive Expressions of Opinion in Context

arXiv.org Artificial Intelligence

Classic information extraction techniques consist in building questions and answers about the facts. Indeed, it is still a challenge to subjective information extraction systems to identify opinions and feelings in context. In sentiment-based NLP tasks, there are few resources to information extraction, above all offensive or hateful opinions in context. To fill this important gap, this short paper provides a new cross-lingual and contextual offensive lexicon, which consists of explicit and implicit offensive and swearing expressions of opinion, which were annotated in two different classes: context dependent and context-independent offensive. In addition, we provide markers to identify hate speech. Annotation approach was evaluated at the expression-level and achieves high human inter-annotator agreement. The provided offensive lexicon is available in Portuguese and English languages.


Transition to Adulthood for Young People with Intellectual or Developmental Disabilities: Emotion Detection and Topic Modeling

arXiv.org Artificial Intelligence

Transition to Adulthood is an essential life stage for many families. The prior research has shown that young people with intellectual or development disabil-ities (IDD) have more challenges than their peers. This study is to explore how to use natural language processing (NLP) methods, especially unsupervised machine learning, to assist psychologists to analyze emotions and sentiments and to use topic modeling to identify common issues and challenges that young people with IDD and their families have. Additionally, the results were compared to those obtained from young people without IDD who were in tran-sition to adulthood. The findings showed that NLP methods can be very useful for psychologists to analyze emotions, conduct cross-case analysis, and sum-marize key topics from conversational data. Our Python code is available at https://github.com/mlaricheva/emotion_topic_modeling.


Find the Funding: Entity Linking with Incomplete Funding Knowledge Bases

arXiv.org Artificial Intelligence

Automatic extraction of funding information from academic articles adds significant value to industry and research communities, such as tracking research outcomes by funding organizations, profiling researchers and universities based on the received funding, and supporting open access policies. Two major challenges of identifying and linking funding entities are: (i) sparse graph structure of the Knowledge Base (KB), which makes the commonly used graph-based entity linking approaches suboptimal for the funding domain, (ii) missing entities in KB, which (unlike recent zero-shot approaches) requires marking entity mentions without KB entries as NIL. We propose an entity linking model that can perform NIL prediction and overcome data scarcity issues in a time and data-efficient manner. Our model builds on a transformer-based mention detection and bi-encoder model to perform entity linking. We show that our model outperforms strong existing baselines.


Gradient Health, Inc on LinkedIn: Data Requirements for FDA

#artificialintelligence

Did you know: Representative Data We'd like to point out some key statistics from our last post on small study sizes. First of all, the question this article is trying to respond is about the prevalence and extent of small study effects in diagnostic imaging. Reach out to us to know how you can have quick access to millions of diverse medical imaging data and avoid data bias: https://lnkd.in/gVwPPXUB


Silicon Valley can't keep track of your data

Washington Post - Technology News

Dina El-Kassaby, a spokeswoman for Meta, Facebook's parent company, said that the deposition did not mean that the company was failing at security or data access issues. "Our systems are sophisticated and it shouldn't be a surprise that no single company engineer can answer every question about where each piece of user information is stored," she said. "We've built one of the most comprehensive privacy programs to oversee data use across our operations and to carefully manage and protect people's data. We have made -- and continue making -- significant investments to meet our privacy commitments and obligations, including extensive data controls."


Automatic Error Analysis for Document-level Information Extraction

arXiv.org Artificial Intelligence

Document-level information extraction (IE) tasks have recently begun to be revisited in earnest using the end-to-end neural network techniques that have been successful on their sentence-level IE counterparts. Evaluation of the approaches, however, has been limited in a number of dimensions. In particular, the precision/recall/F1 scores typically reported provide few insights on the range of errors the models make. We build on the work of Kummerfeld and Klein (2013) to propose a transformation-based framework for automating error analysis in document-level event and (N-ary) relation extraction. We employ our framework to compare two state-of-the-art document-level template-filling approaches on datasets from three domains; and then, to gauge progress in IE since its inception 30 years ago, vs. four systems from the MUC-4 (1992) evaluation.


CommunityLM: Probing Partisan Worldviews from Language Models

arXiv.org Artificial Intelligence

As political attitudes have diverged ideologically in the United States, political speech has diverged lingusitically. The ever-widening polarization between the US political parties is accelerated by an erosion of mutual understanding between them. We aim to make these communities more comprehensible to each other with a framework that probes community-specific responses to the same survey questions using community language models CommunityLM. In our framework we identify committed partisan members for each community on Twitter and fine-tune LMs on the tweets authored by them. We then assess the worldviews of the two groups using prompt-based probing of their corresponding LMs, with prompts that elicit opinions about public figures and groups surveyed by the American National Election Studies (ANES) 2020 Exploratory Testing Survey. We compare the responses generated by the LMs to the ANES survey results, and find a level of alignment that greatly exceeds several baseline methods. Our work aims to show that we can use community LMs to query the worldview of any group of people given a sufficiently large sample of their social media discussions or media diet.


Twitter data unprotected, ex-security chief tells U.S. Congress, as Musk deal approved

The Japan Times

Washington – Twitter whistleblower Peiter Zatko told the U.S. Congress on Tuesday that the platform ignored his security concerns, in testimony that came as company shareholders greenlit Elon Musk's $44 billion takeover deal. Nearly 99% of the votes cast by stock owners endorsed the agreement with Musk to sell him the tech firm for $54.20 per share, Twitter said in a release. This could be due to a conflict with your ad-blocking or security software. Please add japantimes.co.jp and piano.io to your list of allowed sites. If this does not resolve the issue or you are unable to add the domains to your allowlist, please see this support page.


What TikTok and Facebook may track with their in-app browsers

Washington Post - Technology News

Some other iOS social apps, including LinkedIn and Snapchat, also use custom browsers but don't appear to inject similar code, according to Krause's analysis tool, which he made available to the public. Twitter, Reddit and others use Apple's browser, they confirmed, which prevents apps from observing people's activity after they open outside links. A spokeswoman for Twitter said the company switched to Apple's tool in part to protect user privacy.


Twitter's data center knocked out by extreme heat in California

Los Angeles Times

Extreme heat that exhausted California's overworked electric grid on Labor Day had knocked out one of Twitter's main data centers in Sacramento, according to a report. While Twitter avoided a shutdown on Sept. 5 by leaning on its other data centers in Portland, Ore., and Atlanta during the outage to keep its systems running, a company executive warned that if another center were lost, some users would have been unable to access the social media platform, according to an internal memo obtained by CNN. Temperatures in Sacramento on Labor Day broke a daily record of 114 degrees, punching thermometers up to 116 by the afternoon. To power their online services to users, tech companies such as Twitter, Google, or Meta lean on data centers that can demand heavy loads of power and often generate large amounts of heat, requiring cooling systems to keep things running. As climate change continues to heat the planet, Twitter's outage underscores how such extreme weather impacts the online systems that billions of people rely on daily.