Goto

Collaborating Authors

 Oceania


Benchmarking zero-shot and few-shot approaches for tokenization, tagging, and dependency parsing of Tagalog text

arXiv.org Artificial Intelligence

The grammatical analysis of texts in any written language typically involves a number of basic processing tasks, such as tokenization, morphological tagging, and dependency parsing. State-of-the-art systems can achieve high accuracy on these tasks for languages with large datasets, but yield poor results for languages which have little to no annotated data. To address this issue for the Tagalog language, we investigate the use of alternative language resources for creating task-specific models in the absence of dependency-annotated Tagalog data. We also explore the use of word embeddings and data augmentation to improve performance when only a small amount of annotated Tagalog data is available. We show that these zero-shot and few-shot approaches yield substantial improvements on grammatical analysis of both in-domain and out-of-domain Tagalog text compared to state-of-the-art supervised baselines.


Plant species richness prediction from DESIS hyperspectral data: A comparison study on feature extraction procedures and regression models

arXiv.org Artificial Intelligence

The diversity of terrestrial vascular plants plays a key role in maintaining the stability and productivity of ecosystems. Monitoring species compositional diversity across large spatial scales is challenging and time consuming. The advanced spectral and spatial specification of the recently launched DESIS (the DLR Earth Sensing Imaging Spectrometer) instrument provides a unique opportunity to test the potential for monitoring plant species diversity with spaceborne hyperspectral data. This study provides a quantitative assessment on the ability of DESIS hyperspectral data for predicting plant species richness in two different habitat types in southeast Australia. Spectral features were first extracted from the DESIS spectra, then regressed against on-ground estimates of plant species richness, with a two-fold cross validation scheme to assess the predictive performance. We tested and compared the effectiveness of Principal Component Analysis (PCA), Canonical Correlation Analysis (CCA), and Partial Least Squares analysis (PLS) for feature extraction, and Kernel Ridge Regression (KRR), Gaussian Process Regression (GPR), Random Forest Regression (RFR) for species richness prediction. The best prediction results were r=0.76 and RMSE=5.89 for the Southern Tablelands region, and r=0.68 and RMSE=5.95 for the Snowy Mountains region. Relative importance analysis for the DESIS spectral bands showed that the red-edge, red, and blue spectral regions were more important for predicting plant species richness than the green bands and the near-infrared bands beyond red-edge. We also found that the DESIS hyperspectral data performed better than Sentinel-2 multispectral data in the prediction of plant species richness. Our results provide a quantitative reference for future studies exploring the potential of spaceborne hyperspectral data for plant biodiversity mapping.


Sequentially Controlled Text Generation

arXiv.org Artificial Intelligence

While GPT-2 generates sentences that are remarkably human-like, longer documents can ramble and do not follow human-like writing structure. We study the problem of imposing structure on long-range text. We propose a novel controlled text generation task, sequentially controlled text generation, and identify a dataset, NewsDiscourse as a starting point for this task. We develop a sequential controlled text generation pipeline with generation and editing. We test different degrees of structural awareness and show that, in general, more structural awareness results in higher control-accuracy, grammaticality, coherency and topicality, approaching human-level writing performance.


Lifted Reasoning for Combinatorial Counting

Journal of Artificial Intelligence Research

Combinatorics math problems are often used as a benchmark to test human cognitive and logical problem-solving skills. These problems are concerned with counting the number of solutions that exist in a specific scenario that is sketched in natural language. Humans are adept at solving such problems as they can identify commonly occurring structures in the questions for which a closed-form formula exists for computing the answer. These formulas exploit the exchangeability of objects and symmetries to avoid a brute-force enumeration of all possible solutions. Unfortunately, current AI approaches are still unable to solve combinatorial problems in this way. This paper aims to fill this gap by developing novel AI techniques for representing and solving such problems. It makes the following five contributions. First, we identify a class of combinatorics math problems which traditional lifted counting techniques fail to model or solve efficiently. Second, we propose a novel declarative language for this class of problems. Third, we propose novel lifted solving algorithms bridging probabilistic inference techniques and constraint programming. Fourth, we implement them in a lifted solver that solves efficiently the class of problems under investigation. Finally, we evaluate our contributions on a real-world combinatorics math problems dataset and synthetic benchmarks.


Weka ยป ADMIN Magazine

#artificialintelligence

Everyone has probably heard of machine learning, but how exactly does it work? Does it mean that an intelligent machine makes decisions on behalf of humans? You might want to replace the term "intelligent machine" with "efficient algorithm" and add that this algorithm works with data. In doing so, it delivers a view that captures the essence of the data. Simply put, machine learning focuses on building models that learn from existing data and then uses those models to make logical decisions without requiring human intervention.


ChatGPT can create travel itineraries. Should advisors be worried?: Travel Weekly

#artificialintelligence

You have likely heard of ChatGPT, the artificial intelligence chatbot that can create original college essays that don't get flagged by plagiarism-detection software. Of course, it can do many things beyond confounding educators and delighting students. It can, for instance, write computer code. And, I discovered, give travel planning advice. To see how useful its travel suggestions might be, I began by asking what there is to see in Uturoa, the main town on the French Polynesian island of Raiatea.


Using machine learning to improve the toxicity assessment of chemicals

AIHub

Researchers from the University of Amsterdam, together with colleagues at the University of Queensland and the Norwegian Institute for Water Research, have developed a strategy for assessing the toxicity of chemicals using machine learning. The models developed in this study can lead to substantial improvements when compared to conventional'in silico' assessments based on quantitative structure-activity relationship (QSAR) modelling. According to the researchers, the use of machine learning can vastly improve the hazard assessment of molecules, both in the safe-by-design development of new chemicals and in the evaluation of existing chemicals. The importance of the latter is illustrated by the fact that European and US chemical agencies have listed approximately 800,000 chemicals that have been developed over the years but for which there is little to no knowledge about environmental fate or toxicity. Since an experimental assessment of chemical fate and toxicity requires much time, effort, and resources, modelling approaches are already used to predict hazard indicators.


Flutes, synths, a human voice โ€“ how should electric vehicles sound?

The Guardian

Take a walk down any busy street and the noise can hit like a speaker accidentally left on full volume. The growls of engines accelerating when the traffic light turns green, motorbikes vying for position in the traffic, buses whizzing past and the odd rev-head all compete to be heard. The sound generated by the internal combustion engine has shaped urban life for a century, but that is gradually going to change: by 2050, 90% of cars in Australia will be electric. Australia is developing noise standards for electric cars that may follow similar rules set by the UN or US, industry experts say. But what exactly an electric car, and ultimately our cities, will sound like is under the creative control of carmakers.


Machine-Learning Prediction of the Computed Band Gaps of Double Perovskite Materials

arXiv.org Artificial Intelligence

Prediction of the electronic structure of functional materials is essential for the engineering of new devices. Conventional electronic structure prediction methods based on density functional theory (DFT) suffer from not only high computational cost, but also limited accuracy arising from the approximations of the exchange-correlation functional. Surrogate methods based on machine learning have garnered much attention as a viable alternative to bypass these limitations, especially in the prediction of solid-state band gaps, which motivated this research study. Herein, we construct a random forest regression model for band gaps of double perovskite materials, using a dataset of 1306 band gaps computed with the GLLBSC (Gritsenko, van Leeuwen, van Lenthe, and Baerends solid correlation) functional. Among the 20 physical features employed, we find that the bulk modulus, superconductivity temperature, and cation electronegativity exhibit the highest importance scores, consistent with the physics of the underlying electronic structure. Using the top 10 features, a model accuracy of 85.6% with a root mean square error of 0.64 eV is obtained, comparable to previous studies. Our results are significant in the sense that they attest to the potential of machine learning regressions for the rapid screening of promising candidate functional materials.


Artificial intelligence needs regulations that builds public trust in it

#artificialintelligence

To build trust and confidence in the technology, laws should require organisations and governments to use AI in an ethical, safe and responsible manner that protects peoples' privacy. This means companies and the government must be accountable for the decisions their AI systems make. It means AI systems must be transparent and that an organisation can explain how a person's data is being used by the AI system. It means protections must be put in place to help reduce the risk that AI outputs are not biased or discriminatory. It means individuals are notified when AI is used to make a decision that affects their rights. It means there are boundaries on how high-risk AI systems can be used, and it means individuals have an appropriate legal recourse when those boundaries are broken.