Goto

Collaborating Authors

 Africa


Why LLaMa Is A Big Deal

#artificialintelligence

You might have heard about LLaMa or maybe you haven't. In a nutshell, LLaMa is important because it allows you to run large language models (LLM) like GPT-3 on commodity hardware. In many ways, this is a bit like Stable Diffusion, which similarly allowed normal folks to run image generation models on their own hardware with access to the underlying source code. We've discussed why Stable Diffusion matters and even talked about how it works. LLaMa is a transformer language model from Facebook/Meta research, which is a collection of large models from 7 billion to 65 billion parameters trained on publicly available datasets.


YC's latest batch sure was a lot of 'maybe AI can do... this?'

#artificialintelligence

Sitting through hundreds of startups on YC Demo Days, you're not always sure whether you are actually perceiving patterns or if your brain, as coffee battles with monotony, is inventing them in a kind of pareidolia for business plans. This year, though, the theme was pretty obvious: "AI can do that, probably! Certainly today's AI models are more capable than yesterday's, and yesteryear's. But we've seen over and over how these systems demo well but fall down under systematic requirements or as tools with reliable and repeatable results. It's hard not to see this batch as the precursors of a coming wave of AI-powered shovelware.


Scientists detect alien signals coming from 5 nearby stars

#artificialintelligence

Are we alone in the universe? Scientists may have just moved us closer to answering this question. The team – led by researchers from the University of Toronto – has streamlined the search for extraterrestrial life by using a new algorithm to organize the data from their telescopes into categories, in order to distinguish between real signals and interference. This has allowed them to quickly sort through the information and find patterns, through an artificial intelligence process known as machine learning. They discovered eight extraterrestrial signals that seem to have the hallmarks of technology.


WebBrain: Learning to Generate Factually Correct Articles for Queries by Grounding on Large Web Corpus

arXiv.org Artificial Intelligence

In this paper, we introduce a new NLP task -- generating short factual articles with references for queries by mining supporting evidence from the Web. In this task, called WebBrain, the ultimate goal is to generate a fluent, informative, and factually-correct short article (e.g., a Wikipedia article) for a factual query unseen in Wikipedia. To enable experiments on WebBrain, we construct a large-scale dataset WebBrain-Raw by extracting English Wikipedia articles and their crawlable Wikipedia references. WebBrain-Raw is ten times larger than the previous biggest peer dataset, which can greatly benefit the research community. From WebBrain-Raw, we construct two task-specific datasets: WebBrain-R and WebBrain-G, which are used to train in-domain retriever and generator, respectively. Besides, we empirically analyze the performances of the current state-of-the-art NLP techniques on WebBrain and introduce a new framework ReGen, which enhances the generation factualness by improved evidence retrieval and task-specific pre-training for generation. Experiment results show that ReGen outperforms all baselines in both automatic and human evaluations.


Ensemble Modeling for Time Series Forecasting: an Adaptive Robust Optimization Approach

arXiv.org Artificial Intelligence

Accurate time series forecasting is critical for a wide range of problems with temporal data. Ensemble modeling is a well-established technique for leveraging multiple predictive models to increase accuracy and robustness, as the performance of a single predictor can be highly variable due to shifts in the underlying data distribution. This paper proposes a new methodology for building robust ensembles of time series forecasting models. Our approach utilizes Adaptive Robust Optimization (ARO) to construct a linear regression ensemble in which the models' weights can adapt over time. We demonstrate the effectiveness of our method through a series of synthetic experiments and real-world applications, including air pollution management, energy consumption forecasting, and tropical cyclone intensity forecasting. Our results show that our adaptive ensembles outperform the best ensemble member in hindsight by 16-26% in root mean square error and 14-28% in conditional value at risk and improve over competitive ensemble techniques.


Implementation and (Inverse Modified) Error Analysis for implicitly-templated ODE-nets

arXiv.org Artificial Intelligence

We focus on learning unknown dynamics from data using ODE-nets templated on implicit numerical initial value problem solvers. First, we perform Inverse Modified error analysis of the ODE-nets using unrolled implicit schemes for ease of interpretation. It is shown that training an ODE-net using an unrolled implicit scheme returns a close approximation of an Inverse Modified Differential Equation (IMDE). In addition, we establish a theoretical basis for hyper-parameter selection when training such ODE-nets, whereas current strategies usually treat numerical integration of ODE-nets as a black box. We thus formulate an adaptive algorithm which monitors the level of error and adapts the number of (unrolled) implicit solution iterations during the training process, so that the error of the unrolled approximation is less than the current learning loss. This helps accelerate training, while maintaining accuracy. Several numerical experiments are performed to demonstrate the advantages of the proposed algorithm compared to nonadaptive unrollings, and validate the theoretical analysis. We also note that this approach naturally allows for incorporating partially known physical terms in the equations, giving rise to what is termed ``gray box" identification.


Can AI Help Us Save the Planet From Ourselves?

#artificialintelligence

Much of the conversation around artificial intelligence (AI) these days centers on whether it will eventually take your job, how it's trying to compete with humans in creative fields, or how it can be misused, say, as a writing tool. You can probably chalk this one-sidedness up to an all-too-human tendency to be suspicious of new tech that isn't well understood by the mainstream (yet). But AI isn't intrinsically evil or good: It's a tool, a vast technology with enormous potential, and there are myriad ways to implement it beyond the current discourse. One vitally important use case is helping us fight and survive the consequences of climate change. Whether it's mitigating the effects of disasters such as floods and fires more quickly or building a cleaner energy grid, the evidence is mounting that AI has an essential role to play in helping to protect us as the planet reacts to climate change.


Meta's New AI Tool Makes It Easier For Researchers To Analyze Photos

#artificialintelligence

The AI based tool can create "cutouts" or segments of different parts of an image. This comes handy while editing photos or while analyzing imagery for biological or security purposes. These tasks have one thing in common: you need to be able to identify and separate different objects within an image. Traditionally, researchers have had to start from scratch each time they want to analyze a new part of an image. Meta aims to change this laborious process by being the one-stop-shop for researchers and web developers working on such problems.


Decoder-Only or Encoder-Decoder? Interpreting Language Model as a Regularized Encoder-Decoder

arXiv.org Artificial Intelligence

The sequence-to-sequence (seq2seq) task aims at generating the target sequence based on the given input source sequence. Traditionally, most of the seq2seq task is resolved by the Encoder-Decoder framework which requires an encoder to encode the source sequence and a decoder to generate the target text. Recently, a bunch of new approaches have emerged that apply decoder-only language models directly to the seq2seq task. Despite the significant advancements in applying language models to the seq2seq task, there is still a lack of thorough analysis on the effectiveness of the decoder-only language model architecture. This paper aims to address this gap by conducting a detailed comparison between the encoder-decoder architecture and the decoder-only language model framework through the analysis of a regularized encoder-decoder structure. This structure is designed to replicate all behaviors in the classical decoder-only language model but has an encoder and a decoder making it easier to be compared with the classical encoder-decoder structure. Based on the analysis, we unveil the attention degeneration problem in the language model, namely, as the generation step number grows, less and less attention is focused on the source sequence. To give a quantitative understanding of this problem, we conduct a theoretical sensitivity analysis of the attention output with respect to the source input. Grounded on our analysis, we propose a novel partial attention language model to solve the attention degeneration problem. Experimental results on machine translation, summarization, and data-to-text generation tasks support our analysis and demonstrate the effectiveness of our proposed model.


Pump It Up: Predict Water Pump Status using Attentive Tabular Learning

arXiv.org Artificial Intelligence

Water crisis is a crucial concern around the globe. Appropriate and timely maintenance of water pumps in drought-hit countries is vital for communities relying on the well. In this paper, we analyze and apply a sequential attentive deep neural architecture, TabNet, for predicting water pump repair status in Tanzania. The model combines the valuable benefits of tree-based algorithms and neural networks, enabling end-to-end training, model interpretability, sparse feature selection, and efficient learning on tabular data. Finally, we compare the performance of TabNet with popular gradient tree-boosting algorithms like XGBoost, LightGBM,CatBoost, and demonstrate how we can further uplift the performance by choosing focal loss as the objective function while training on imbalanced data.