Genre
Thousands of fMRI brain studies in doubt due to software flaws
The discovery of major software flaws could render thousands of fMRI brain studies inaccurate. The use of fMRI is a common method for scanning the brain in neuroscience and psychology experiments. To make sense of the data produced, researchers sometimes use a technique called spatial autocorrelation to identify areas of the brain that appear to "light up" during particular tasks or experiences. But some software flaws in the popular fMRI data analysis packages SPM, FSL and AFNI meant this technique routinely produced false positives, resulting in errors 50 per cent of the time or more. Anders Eklund and Hans Knutsson at Linkรถping University in Sweden and Thomas Nichols at the University of Warwick, UK, calculated this by analysing brain data from a collaborative open fMRI project called 1000 Functional Connectomes.
Understanding the impact of AI
Coding will join this list in time, however, where it differs wildly from the afore mentioned examples is it is unlikely to be lovingly preserved for future generations to admire, fiddle with or better still, reactivate. Its essence will not be reified for one specific reason โ it can't be touched and humans value tactility. We touch immediately, both inside and outside the womb. Today, we find ourselves at a pivotal moment in our existence and about to experience an exponential period of rapid technological growth the likes of which is quite probably beyond our comprehension and at a base level, will have serious implications for coding. We rather arrogantly think that because we have a good grasp of our own technological advancement so far, we can somehow predict the mass cultural and behavioural shift about to happen as we question our own skills in the world. Us techies hold on to the notion that we are the masters of code, and we will be forever commanding line by line, the computers to do our bidding.
Artificial Intelligence Latest News & Updates: How Can Machine Learning Play A Significant Role In Autism Diagnosis And Intervention?
Major landmarks around the world are Lighting It Up Blue on April 1 and 2 to raise awareness about Autism Spectrum Disorders (ASD) for World Autism Awareness Day at Forte Sangallo on April 02, 2016 in Nettuno, Italy. In recent months, artificial intelligence (AI) has been making its presence known in different fields of sciences. In fact, AI is deemed as a valuable asset in precision medicine. But now, a team of researchers is exploring the possibilities if machine learning could play a vital part in autism screening, diagnostics and intervention. Before delving deeper into the latest research on the importance of artificial intelligence in autism screening and diagnostics, let's first define the two most relevant subjects on the study - autism and machine learning. According to Autism Speaks, autism refers to the "general term used for group of complex disorders of brain development," which are marked by social interaction, verbal and nonverbal communication difficulties, as well as repetitive behaviors.
On the Application of Support Vector Machines to the Prediction of Propagation Losses at 169 MHz for Smart Metering Applications
Uccellari, Martino, Facchini, Francesca, Sola, Matteo, Sirignano, Emilio, Vitetta, Giorgio M., Barbieri, Andrea, Tondelli, Stefano
Recently, the need of deploying new wireless networks for smart gas metering has raised the problem of radio planning in the169 MHz band. Unluckily, software tools commonly adopted for radio planning in cellular communication systems cannot be employed to solve this problem because of the substantially lower transmission frequencies characterizing this application. In this manuscript a novel data-centric solution, based on the use of support vector machine techniques for classification and regression, is proposed. Our method requires the availability of a limited set of received signal strength measurements and the knowledge of a three-dimensional map of the propagation environment of interest, and generates both an estimate of the coverage area and a prediction of the field strength within it. Numerical results referring to different Italian villages and cities evidence that our method is able to achieve good accuracy at the price of an acceptable computational cost and of a limited effort for the acquisition of measurements in the considered environments.
A Batch, Off-Policy, Actor-Critic Algorithm for Optimizing the Average Reward
Murphy, S. A., Deng, Y., Laber, E. B., Maei, H. R., Sutton, R. S., Witkiewitz, K.
We develop an off-policy actor-critic algorithm for learning an optimal policy from a training set composed of data from multiple individuals. This algorithm is developed with a view toward its use in mobile health. In the behavioral health communities there is increasing interest in, and use of, mobile devices to deliver treatments that target behavior change. Mobile devices can be used to provide treatment when, where, and in the amount desired (Litvin et al., 2013; Kumar et al., 2013). Increasingly scientists are looking to passive sensing (wearable devices, GPS, activity on the smartphone) and self-report of internal states to individualize the intervention to the person in terms of when, how and where to deliver treatment.
Geometric Mean Metric Learning
Zadeh, Pourya Habib, Hosseini, Reshad, Sra, Suvrit
We revisit the task of learning a Euclidean metric from data. We approach this problem from first principles and formulate it as a surprisingly simple optimization problem. Indeed, our formulation even admits a closed form solution. This solution possesses several very attractive properties: (i) an innate geometric appeal through the Riemannian geometry of positive definite matrices; (ii) ease of interpretability; and (iii) computational speed several orders of magnitude faster than the widely used LMNN and ITML methods. Furthermore, on standard benchmark datasets, our closed-form solution consistently attains higher classification accuracy.
Graphical Model Sketch
Kveton, Branislav, Bui, Hung, Ghavamzadeh, Mohammad, Theocharous, Georgios, Muthukrishnan, S., Sun, Siqi
Structured high-cardinality data arises in many domains, and poses a major challenge for both modeling and inference. Graphical models are a popular approach to modeling structured data but they are unsuitable for high-cardinality variables. The count-min (CM) sketch is a popular approach to estimating probabilities in high-cardinality data but it does not scale well beyond a few variables. In this work, we bring together the ideas of graphical models and count sketches; and propose and analyze several approaches to estimating probabilities in structured high-cardinality streams of data. The key idea of our approximations is to use the structure of a graphical model and approximately estimate its factors by "sketches", which hash high-cardinality variables using random projections. Our approximations are computationally efficient and their space complexity is independent of the cardinality of variables. Our error bounds are multiplicative and significantly improve upon those of the CM sketch, a state-of-the-art approach to estimating probabilities in streams. We evaluate our approximations on synthetic and real-world problems, and report an order of magnitude improvements over the CM sketch.
On the use of Harrell's C for clinical risk prediction via random survival forests
Schmid, Matthias, Wright, Marvin, Ziegler, Andreas
Random survival forests (RSF) are a powerful method for risk prediction of right-censored outcomes in biomedical research. RSF use the log-rank split criterion to form an ensemble of survival trees. The most common approach to evaluate the prediction accuracy of a RSF model is Harrell's concordance index for survival data ('C index'). Conceptually, this strategy implies that the split criterion in RSF is different from the evaluation criterion of interest. This discrepancy can be overcome by using Harrell's C for both node splitting and evaluation. We compare the difference between the two split criteria analytically and in simulation studies with respect to the preference of more unbalanced splits, termed end-cut preference (ECP). Specifically, we show that the log-rank statistic has a stronger ECP compared to the C index. In simulation studies and with the help of two medical data sets we demonstrate that the accuracy of RSF predictions, as measured by Harrell's C, can be improved if the log-rank statistic is replaced by the C index for node splitting. This is especially true in situations where the censoring rate or the fraction of informative continuous predictor variables is high. Conversely, log-rank splitting is preferable in noisy scenarios. Both C-based and log-rank splitting are implemented in the R~package ranger. We recommend Harrell's C as split criterion for use in smaller scale clinical studies and the log-rank split criterion for use in large-scale 'omics' studies.
Machine Learning Meta-analysis of Large Metagenomic Datasets: Tools and Biological Insights
Shotgun metagenomic analysis of the human associated microbiome provides a rich set of microbial features for prediction and biomarker discovery in the context of human diseases and health conditions. However, the use of such high-resolution microbial features presents new challenges, and validated computational tools for learning tasks are lacking. Moreover, classification rules have scarcely been validated in independent studies, posing questions about the generality and generalization of disease-predictive models across cohorts. In this paper, we comprehensively assess approaches to metagenomics-based prediction tasks and for quantitative assessment of the strength of potential microbiome-phenotype associations. We develop a computational framework for prediction tasks using quantitative microbiome profiles, including species-level relative abundances and presence of strain-specific markers.
How a Technical Co-founder Spends his Time: Minute-by-minute Data for a Year
I'm co-founder and CTO at Overleaf, a successful SaaS startup based in London. From August 2014 to December 2015, I manually tracked all of my work time, minute-by-minute, and analysed the data in R. Like most people who track their time, my goal was to improve my productivity. It gave me data to answer questions about whether I was spending too much or too little time on particular activities, for example user support or client projects. The data showed that my intuition on these questions was often wrong. There were also some less tangible benefits. It was reassuring on a Friday to have an answer to that usually rhetorical question, "where did this week go?" I feel like it also reduced context switching: if I stopped what I was doing to answer an chat message or email, I had to take the time to record it in my time tracker. I think this added friction was a win for overall productivity, perhaps paradoxically. This post documents the (simple) system I built to record my time, how I analysed the data, and the results.