Africa
World's first AI health app in Swahili launches to tackle doctor shortages
An innovative chat-bot that helps patients and doctors diagnose diseases ranging from malaria to diabetes has become the first health app to launch in Swahili. Developed by Ada Health, the app relies on artificial intelligence, large medical databases and personalised responses to assess an individual's symptoms, suggest a cause and recommend the next stage of treatment. The smartphone chat-bot is already used by roughly eight million people in more than 130 countries across the globe – published in languages including English, French and Spanish. But it has now become the first AI health application to launch in Swahili, a language spoken by almost 100 million people across East Africa – predominantly in Tanzania, Uganda and Kenya. According to Hila Azadzoy, the managing director of Ada's global health initiative, the expansion will help tackle a shortage of doctors and nurses in the region, where countries have fewer than one physician per 1,000 people on average.
Robotic Processing Automation, Hello November
Sign in to report inappropriate content. It's been some time since there's been a video on my Vlog Channel. Good to be back, on this video I share some quick snippets of my recent trip to #NewYorkCity. I Dive into the great work we have been doing at Hashtag South Africa and welcoming you to Robotic Processing Automation, and Artifical Intelligence solutions we are now providing to our customers. As Usual, I'm recapping Global Goals 2030 and how everyone around the world is working together.
Artificial Intelligence Bias, Russia, Fentanyl: RAND Weekly Recap
This week, we discuss what to do about bias in algorithms; Russia's limits in the Middle East; learning from other countries' experiences with fentanyl; what protests could mean for democracy in the Middle East; how cities can help U.S. diplomacy; and helping U.S. Army special operations forces assess their missions. Earlier this month, a controversy about gender bias in the Apple Card algorithm lit up social media; an outraged tech executive posted about how his credit line was 20 times higher than his wife's, even though the two share all assets. According to RAND's Osonde Osoba, problems like this may become more common as artificial intelligence is used in more kinds of decisionmaking. It's not always possible to pinpoint how a complex algorithm led to a bad outcome, he says. But there are ways for companies to audit algorithms for sexist, racist, biased behaviors.
Machine Learning – It all Boils Down to the Training Data Vinod Sharma's Blog
Training Data – There is a famous punch line about Data. "Data is not good enough if it's not a quality data". If your data model is not working or performing as expected blame your data and data source. Instead of struggling to find an opportunity for performance tuning look for improving data quality fed into the model. AILabPage defines machine learning as "A focal point where business, data, experience meets emerging technology and decides to work together".
AI has a bias problem. Barring African experts from a conference in Canada won't help
London (CNN Business)Some of the leading artificial intelligence experts from Africa and South America have been denied visas to attend a major industry conference in Canada, dealing a setback to efforts to prevent bias from taking root in the new technology. Conference organizers say Canadian immigration authorities have denied visas to two dozen academics from countries such as Nigeria and Brazil, preventing them from attending the event next month in Vancouver. Katherine Heller, a professor who serves as co-chair of diversity and inclusion at the Neural Information Processing Systems conference, said organizers "are trying extremely hard" to have the visa denials overturned. "It is very significant for the field of AI that all voices be heard," she said. The problem of algorithmic bias in data science has become more pronounced, and there's mounting evidence that AI-powered algorithms display bias against women and some racial groups.
Fast and Scalable Estimator for Sparse and Unit-Rank Higher-Order Regression Models
Because tensor data appear more and more frequently in various scientific researches and real-world applications, analyzing the relationship between tensor features and the univariate outcome becomes an elementary task in many fields. To solve this task, we propose \underline{Fa}st \underline{S}parse \underline{T}ensor \underline{R}egression model (FasTR) based on so-called unit-rank CANDECOMP/PARAFAC decomposition. FasTR first decomposes the tensor coefficient into component vectors and then estimates each vector with $\ell_1$ regularized regression. Because of the independence of component vectors, FasTR is able to solve in a parallel way and the time complexity is proved to be superior to previous models. We evaluate the performance of FasTR on several simulated datasets and a real-world fMRI dataset. Experiment results show that, compared with four baseline models, in every case, FasTR can compute a better solution within less time.
On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks
Şimşekli, Umut, Gürbüzbalaban, Mert, Nguyen, Thanh Huy, Richard, Gaël, Sagun, Levent
The gradient noise (GN) in the stochastic gradient descent (SGD) algorithm is often considered to be Gaussian in the large data regime by assuming that the \emph{classical} central limit theorem (CLT) kicks in. This assumption is often made for mathematical convenience, since it enables SGD to be analyzed as a stochastic differential equation (SDE) driven by a Brownian motion. We argue that the Gaussianity assumption might fail to hold in deep learning settings and hence render the Brownian motion-based analyses inappropriate. Inspired by non-Gaussian natural phenomena, we consider the GN in a more general context and invoke the \emph{generalized} CLT, which suggests that the GN converges to a \emph{heavy-tailed} $\alpha$-stable random vector, where \emph{tail-index} $\alpha$ determines the heavy-tailedness of the distribution. Accordingly, we propose to analyze SGD as a discretization of an SDE driven by a L\'{e}vy motion. Such SDEs can incur `jumps', which force the SDE and its discretization \emph{transition} from narrow minima to wider minima, as proven by existing metastability theory and the extensions that we proved recently. In this study, under the $\alpha$-stable GN assumption, we further establish an explicit connection between the convergence rate of SGD to a local minimum and the tail-index $\alpha$. To validate the $\alpha$-stable assumption, we conduct experiments on common deep learning scenarios and show that in all settings, the GN is highly non-Gaussian and admits heavy-tails. We investigate the tail behavior in varying network architectures and sizes, loss functions, and datasets. Our results open up a different perspective and shed more light on the belief that SGD prefers wide minima.
Short Term Prediction of Parking Area states Using Real Time Data and Machine Learning Techniques
Provoost, Jesper, Wismans, Luc, Van der Drift, Sander, Kamilaris, Andreas, Van Keulen, Maurice
Public road authorities and private mobility service providers need information derived from the current and predicted traffic states to act upon the daily urban system and its spatial and temporal dynamics. In this research, a real-time parking area state (occupancy, in- and outflux) prediction model (up to 60 minutes ahead) has been developed using publicly available historic and real time data sources. Based on a case study in a real-life scenario in the city of Arnhem, a Neural Network-based approach outperforms a Random Forest-based one on all assessed performance measures, although the differences are small. Both are outperforming a naive seasonal random walk model. Although the performance degrades with increasing prediction horizon, the model shows a performance gain of over 150% at a prediction horizon of 60 minutes compared with the naive model. Furthermore, it is shown that predicting the in- and outflux is a far more difficult task (i.e. performance gains of 30%) which needs more training data, not based exclusively on occupancy rate. However, the performance of predicting in- and outflux is less sensitive to the prediction horizon. In addition, it is shown that real-time information of current occupancy rate is the independent variable with the highest contribution to the performance, although time, traffic flow and weather variables also deliver a significant contribution. During real-time deployment, the model performs three times better than the naive model on average. As a result, it can provide valuable information for proactive traffic management as well as mobility service providers.
Detecting anthropogenic cloud perturbations with deep learning
Watson-Parris, Duncan, Sutherland, Samuel, Christensen, Matthew, Caterini, Anthony, Sejdinovic, Dino, Stier, Philip
One of the most pressing questions in climate science is that of the effect of anthropogenic aerosol on the Earth's energy balance. Aerosols provide the `seeds' on which cloud droplets form, and changes in the amount of aerosol available to a cloud can change its brightness and other physical properties such as optical thickness and spatial extent. Clouds play a critical role in moderating global temperatures and small perturbations can lead to significant amounts of cooling or warming. Uncertainty in this effect is so large it is not currently known if it is negligible, or provides a large enough cooling to largely negate present-day warming by CO2. This work uses deep convolutional neural networks to look for two particular perturbations in clouds due to anthropogenic aerosol and assess their properties and prevalence, providing valuable insights into their climatic effects.
Sparse and Low-Rank Tensor Regression via Parallel Proximal Method
Motivated by applications in various scientific fields having demand of predicting relationship between higher-order (tensor) feature and univariate response, we propose a \underline{S}parse and \underline{L}ow-rank \underline{T}ensor \underline{R}egression model (SLTR). This model enforces sparsity and low-rankness of the tensor coefficient by directly applying $\ell_1$ norm and tensor nuclear norm on it respectively, such that (1) the structural information of tensor is preserved and (2) the data interpretation is convenient. To make the solving procedure scalable and efficient, SLTR makes use of the proximal gradient method to optimize two norm regularizers, which can be easily implemented parallelly. Additionally, a tighter convergence rate is proved over three-order tensor data. We evaluate SLTR on several simulated datasets and one fMRI dataset. Experiment results show that, compared with previous models, SLTR is able to obtain a solution no worse than others with much less time cost.