Goto

Collaborating Authors

 Education


How to Find Actionable Static Analysis Warnings: A Case Study with FindBugs

arXiv.org Artificial Intelligence

Automatically generated static code warnings suffer from a large number of false alarms. Hence, developers only take action on a small percent of those warnings. To better predict which static code warnings should not be ignored, we suggest that analysts need to look deeper into their algorithms to find choices that better improve the particulars of their specific problem. Specifically, we show here that effective predictors of such warnings can be created by methods that locally adjust the decision boundary (between actionable warnings and others). These methods yield a new high water-mark for recognizing actionable static code warnings. For eight open-source Java projects (cassandra, jmeter, commons, lucene-solr, maven, ant, tomcat, derby) we achieve perfect test results on 4/8 datasets and, overall, a median AUC (area under the true negatives, true positives curve) of 92%.


STRUDEL: Structured Dialogue Summarization for Dialogue Comprehension

arXiv.org Artificial Intelligence

Abstractive dialogue summarization has long been viewed as an important standalone task in natural language processing, but no previous work has explored the possibility of whether abstractive dialogue summarization can also be used as a means to boost an NLP system's performance on other important dialogue comprehension tasks. In this paper, we propose a novel type of dialogue summarization task - STRUctured DiaLoguE Summarization - that can help pre-trained language models to better understand dialogues and improve their performance on important dialogue comprehension tasks. We further collect human annotations of STRUDEL summaries over 400 dialogues and introduce a new STRUDEL dialogue comprehension modeling framework that integrates STRUDEL into a graph-neural-network-based dialogue reasoning module over transformer encoder language models to improve their dialogue comprehension abilities. In our empirical experiments on two important downstream dialogue comprehension tasks - dialogue question answering and dialogue response prediction - we show that our STRUDEL dialogue comprehension model can significantly improve the dialogue comprehension performance of transformer encoder language models.


Image Classification with Small Datasets: Overview and Benchmark

arXiv.org Artificial Intelligence

Image classification with small datasets has been an active research area in the recent past. However, as research in this scope is still in its infancy, two key ingredients are missing for ensuring reliable and truthful progress: a systematic and extensive overview of the state of the art, and a common benchmark to allow for objective comparisons between published methods. This article addresses both issues. First, we systematically organize and connect past studies to consolidate a community that is currently fragmented and scattered. Second, we propose a common benchmark that allows for an objective comparison of approaches. It consists of five datasets spanning various domains (e.g., natural images, medical imagery, satellite data) and data types (RGB, grayscale, multispectral). We use this benchmark to re-evaluate the standard cross-entropy baseline and ten existing methods published between 2017 and 2021 at renowned venues. Surprisingly, we find that thorough hyper-parameter tuning on held-out validation data results in a highly competitive baseline and highlights a stunted growth of performance over the years. Indeed, only a single specialized method dating back to 2019 clearly wins our benchmark and outperforms the baseline classifier.


Utilizing Priming to Identify Optimal Class Ordering to Alleviate Catastrophic Forgetting

arXiv.org Artificial Intelligence

In order for artificial neural networks to begin accurately mimicking biological ones, they must be able to adapt to new exigencies without forgetting what they have learned from previous training. Lifelong learning approaches to artificial neural networks attempt to strive towards this goal, yet have not progressed far enough to be realistically deployed for natural language processing tasks. The proverbial roadblock of catastrophic forgetting still gate-keeps researchers from an adequate lifelong learning model. While efforts are being made to quell catastrophic forgetting, there is a lack of research that looks into the importance of class ordering when training on new classes for incremental learning. This is surprising as the ordering of "classes" that humans learn is heavily monitored and incredibly important. While heuristics to develop an ideal class order have been researched, this paper examines class ordering as it relates to priming as a scheme for incremental class learning. By examining the connections between various methods of priming found in humans and how those are mimicked yet remain unexplained in life-long machine learning, this paper provides a better understanding of the similarities between our biological systems and the synthetic systems while simultaneously improving current practices to combat catastrophic forgetting. Through the merging of psychological priming practices with class ordering, this paper is able to identify a generalizable method for class ordering in NLP incremental learning tasks that consistently outperforms random class ordering.


The choice of scaling technique matters for classification performance

arXiv.org Artificial Intelligence

Dataset scaling, also known as normalization, is an essential preprocessing step in a machine learning pipeline. It is aimed at adjusting attributes scales in a way that they all vary within the same range. This transformation is known to improve the performance of classification models, but there are several scaling techniques to choose from, and this choice is not generally done carefully. In this paper, we execute a broad experiment comparing the impact of 5 scaling techniques on the performances of 20 classification algorithms among monolithic and ensemble models, applying them to 82 publicly available datasets with varying imbalance ratios. Results show that the choice of scaling technique matters for classification performance, and the performance difference between the best and the worst scaling technique is relevant and statistically significant in most cases. They also indicate that choosing an inadequate technique can be more detrimental to classification performance than not scaling the data at all. We also show how the performance variation of an ensemble model, considering different scaling techniques, tends to be dictated by that of its base model. Finally, we discuss the relationship between a model's sensitivity to the choice of scaling technique and its performance and provide insights into its applicability on different model deployment scenarios. Full results and source code for the experiments in this paper are available in a GitHub repository.\footnote{https://github.com/amorimlb/scaling\_matters}


Machine Learning Impact in 2022 โ€“ The Official Blog of BigML.com

#artificialintelligence

We are about to wrap up 2022, a year that brought plenty of Machine Learning projects, events, education opportunities, and many groundbreaking Machine Learning applications developed by ML practitioners around the world. The challenges and business needs of our customers continue to fuel our passion to bring to life the robust and innovative Machine Learning solutions they deserve. In this blog post, we put together the highlights of 2022 covering Machine Learning's lasting impact on a vast number of industries and businesses, BigML's new additions and enhancements to our pioneering Machine Learning software platform, our live and virtual events, education initiative updates, and much more! None of the numbers listed above and the activities described on this blog post would be possible without our customers, partners, followers, and certified practitioners. That's why this blog post is dedicated to all of you.


AI can now write like a human. Some teachers are worried.

#artificialintelligence

"The 360" shows you diverse perspectives on the day's top stories and debates. Artificial intelligence has advanced at an extraordinary pace over the past few years. Today, these incredibly complex algorithms are capable of creating award-winning art, penning scripts that can be turned into real films and -- in the latest step that has dazzled people in the tech and media industries -- mimic writing at a level so convincing that it's impossible to tell whether the words were put together by a human or a machine. A few weeks ago, the research company OpenAI released ChatGPT, a language model that can construct remarkably well-structured arguments based on simple prompts provided by a user. The system -- which uses a massive repository of online text to predict what words should come next -- is able to create new stories in the style of famous writers, write news articles about itself and produce essays that could easily receive a passing grade in most English classes.


An AI-based platform to enhance and personalize e-learning

#artificialintelligence

Researchers at Universidad Autรณnoma de Madrid have recently created an innovative, AI-powered platform that could enhance remote learning, allowing educators to securely monitor students and verify that they are attending compulsory online classes or exams. An initial prototype of this platform, called Demo-edBB, is set to be presented at the AAAI-23 Conference on Artificial Intelligence in February 2022, in Washington, and a version of the paper is available on the arXiv preprint server. "Our investigation group, the BiDA-Lab at Universidad Autรณnoma de Madrid, has substantial experience with biometric signals and systems, behavior analysis and AI applications, with over 300 hundred published papers in last two decades," Roberto Daza Garcia, one of the researchers who carried out the study, told TechXplore. "Over the past few years, virtual education has grown significantly, becoming the main foundation of one on the most important educational institutions and generating new valuable opportunities for learning. Our group has thus recently been working on new technologies for e-learning, ultimately leading to the development of a platform that combines biometric and behavior analysis tools."


I'm Convinced My Child's Teacher Has It Out for Her

Slate

What is the best way to handle a high school teacher who just seems to--no matter what my child does--view her as a B student? It's her French class, and the class has many exams that are subjective. For instance, rubrics for oral presentations and written work differentiates between A's and B's by work being "very organized" and "organized," or "often uses complex sentences" and just "uses complex sentences." My child asks for feedback on her work, and she's (noticeably to the teacher) making an effort in class. The feedback she gets is as elusive as the rubrics--not very detailed, and it's not clear how to do "more" of what the teacher describes, since my child is already doing it. My child has always been a straight-A student, and she's incredibly stressed by this class, which she's currently getting a B in.


The Real A.I. In College Admission

#artificialintelligence

I know what you're thinking. "Another article about ChatGPT (Generative Pre-trained Transformer), the artificial intelligence wonder-bot from OpenAI, and how it is going to revolutionize society, work, education, and more." Perhaps it will, but that is not this article. The rise of this extraordinary technology, rather than muddling it, makes it clearer than ever what constitutes authenticity. If you or your child are applying to college, you have undoubtedly heard an admission officer talk about authenticity.