Deep Learning
Accurate Cup-to-Disc Ratio Measurement with Tight Bounding Box Supervision in Fundus Photography
The cup-to-disc ratio (CDR) is one of the most significant indicator for glaucoma diagnosis. Different from the use of costly fully supervised learning formulation with pixel-wise annotations in the literature, this study investigates the feasibility of accurate CDR measurement in fundus images using only tight bounding box supervision. For this purpose, we develop a two-task network for accurate CDR measurement, one for weakly supervised image segmentation, and the other for bounding-box regression. The weakly supervised image segmentation task is implemented based on generalized multiple instance learning formulation and smooth maximum approximation, and the bounding-box regression task outputs class-specific bounding box prediction in a single scale at the original image resolution. To get accurate bounding box prediction, a class-specific bounding-box normalizer and an expected intersection-over-union are proposed. In the experiments, the proposed approach was evaluated by a testing set with 1200 images using CDR error and F1 score for CDR measurement and dice coefficient for image segmentation. A grader study was conducted to compare the performance of the proposed approach with those of individual graders. The results demonstrate that the proposed approach outperforms the state-of-the-art performance obtained from the fully supervised image segmentation (FSIS) approach using pixel-wise annotation for CDR measurement, which is also better than those of individual graders. It also gets performance close to the state-of-the-art obtained from FSIS for optic cup and disc segmentation, similar to those of individual graders. The codes are available at \url{https://github.com/wangjuan313/CDRNet}.
PL-EESR: Perceptual Loss Based END-TO-END Robust Speaker Representation Extraction
Ma, Yi, Lee, Kong Aik, Hautamaki, Ville, Li, Haizhou
Speech enhancement aims to improve the perceptual quality of the speech signal by suppression of the background noise. However, excessive suppression may lead to speech distortion and speaker information loss, which degrades the performance of speaker embedding extraction. To alleviate this problem, we propose an end-to-end deep learning framework, dubbed PL-EESR, for robust speaker representation extraction. This framework is optimized based on the feedback of the speaker identification task and the high-level perceptual deviation between the raw speech signal and its noisy version. We conducted speaker verification tasks in both noisy and clean environment respectively to evaluate our system. Compared to the baseline, our method shows better performance in both clean and noisy environments, which means our method can not only enhance the speaker relative information but also avoid adding distortions.
Exploration of AI-Oriented Power System Transient Stability Simulations
Xiao, Tannan, Chen, Ying, Wang, Jianquan, Huang, Shaowei, Tong, Weilin, He, Tirui
Artificial Intelligence (AI) has made significant progress in the past 5 years and is playing a more and more important role in power system analysis and control. It is foreseeable that the future power system transient stability simulations will be deeply integrated with AI. However, the existing power system dynamic simulation tools are not AI-friendly enough. In this paper, a general design of an AI-oriented power system transient stability simulator is proposed. It is a parallel simulator with a flexible application programming interface so that the simulator has rapid simulation speed, neural network supportability, and network topology accessibility. A prototype of this design is implemented and made public based on our previously realized simulator. Tests of this AI-oriented simulator are carried out under multiple scenarios, which proves that the design and implementation of the simulator are reasonable, AI-friendly, and highly efficient.
Marginally calibrated response distributions for end-to-end learning in autonomous driving
End-to-end learners for autonomous driving are deep neural networks that predict the instantaneous steering angle directly from images of the ahead-lying street. These learners must provide reliable uncertainty estimates for their predictions in order to meet safety requirements and initiate a switch to manual control in areas of high uncertainty. Yet end-to-end learners typically only deliver point predictions, since distributional predictions are associated with large increases in training time or additional computational resources during prediction. To address this shortcoming we investigate efficient and scalable approximate inference for the implicit copula neural linear model of Klein, Nott and Smith (2021) in order to quantify uncertainty for the predictions of end-to-end learners. The result are densities for the steering angle that are marginally calibrated, i.e.~the average of the estimated densities equals the empirical distribution of steering angles. To ensure the scalability to large $n$ regimes, we develop efficient estimation based on variational inference as a fast alternative to computationally intensive, exact inference via Hamiltonian Monte Carlo. We demonstrate the accuracy and speed of the variational approach in comparison to Hamiltonian Monte Carlo on two end-to-end learners trained for highway driving using the comma2k19 data set. The implicit copula neural linear model delivers accurate calibration, high-quality prediction intervals and allows to identify overconfident learners. Our approach also contributes to the explainability of black-box end-to-end learners, since predictive densities can be used to understand which steering actions the end-to-end learner sees as valid.
Hierarchical Gaussian Process Models for Regression Discontinuity/Kink under Sharp and Fuzzy Designs
We propose nonparametric Bayesian estimators for causal inference exploiting Regression Discontinuity/Kink (RD/RK) under sharp and fuzzy designs. Our estimators are based on Gaussian Process (GP) regression and classification. The GP methods are powerful probabilistic modeling approaches that are advantageous in terms of derivative estimation and uncertainty qualification, facilitating RK estimation and inference of RD/RK models. These estimators are extended to hierarchical GP models with an intermediate Bayesian neural network layer and can be characterized as hybrid deep learning models. Monte Carlo simulations show that our estimators perform similarly and often better than competing estimators in terms of precision, coverage and interval length. The hierarchical GP models improve upon one-layer GP models substantially. An empirical application of the proposed estimators is provided.
A Cascaded Deep Learning-Based Artificial Intelligence Algorithm for Automated Lesion Detection and Classification on Biparametric Prostate Magnetic Resonance Imaging - Docwire News
RATIONALE AND OBJECTIVES: Prostate MRI improves detection of clinically significant prostate cancer; however, its diagnostic performance has wide variation. Artificial intelligence (AI) has the potential to assist radiologists in the detection and classification of prostatic lesions. Herein, we aimed to develop and test a cascaded deep learning detection and classification system trained on biparametric prostate MRI using PI-RADS for assisting radiologists during prostate MRI read out. MATERIALS AND METHODS: T2-weighted, diffusion-weighted (ADC maps, high b value DWI) MRI scans obtained at 3 Tesla from two institutions (n 1043 in-house and n 347 Prostate-X, respectively) acquired between 2015 to 2019 were used for model training, validation, testing. All scans were retrospectively reevaluated by one radiologist.
Yet Another GPT-3 Competitor Emerges - Robot Writers AI
AI21 is the latest company to release a competitor to GPT-3 -- currently the gold standard of supercomputer-powered, AI auto-text generators. Dubbed Jurassic-1, AI21's alternative to GPT-3 weighs in at 178 billion parameters and is available in a limited free trial to everyone. It's also billed as the largest English language AI auto-text generator on the market. This article in VentureBeat offers the results of a test-drive of AI21's auto-text generator. VentureBeat's verdict: Jurassic-1 is a promising alternative to GPT-3, with some limitations, according VentureBeat writer Abhishek Iyer.
Global Big Data Conference
Apple has been slowly but surely creating a name for itself in the low-code/no-code movement. This July, the Cupertino-based company announced the launch of Trinity AI, a no-code platform for complex spatial datasets. Trinity enables machine learning researchers and non-AI devs to tailor complex spatiotemporal datasets to fit deep learning models. Back in 2019, Apple revealed SwiftUI, a programming language that required much less coding than the Swift language. With the release of Trinity, Apple doubles down on its effort to significantly lower the threshold for non-devs and non-ML devs.
Machine Learning, Deep Learning, and Artificial Intelligence: a simple explanation
Today we hear about Artificial Intelligence everywhere. At work, when we look for software that can learn autonomously. Yet, when an expert presents us with a new solution equipped with AI, using terms such as Machine Learning, Deep Learning, and neural networks, we start to have doubts about the real understanding of the topic.
Deep Learning -- A brief Introduction
Gradient descent is a iterative optimization algorithm in Linear Regression to find a local minimum of a differential function. The gradient descent algorithm will choose a random initialization at the beginning for slopes and intercept and calculate the Error term. During every iteration it will back check the error term and back propagate and update the Slope and Intercept values until minimizing the error. Learning rate represents the step size and the Partial derivate represent the direction towards it should move. Always recommended to choose a smaller Learning rate so that it will not miss the local minimum.