Regression
Solar Power Time Series Forecasting Utilising Wavelet Coefficients
Almaghrabi, Sarah, Rana, Mashud, Hamilton, Margaret, Rahaman, Mohammad Saiedur
Accurate and reliable prediction of Photovoltaic (PV) power output is critical to electricity grid stability and power dispatching capabilities. However, Photovoltaic (PV) power generation is highly volatile and unstable due to different reasons. The Wavelet Transform (WT) has been utilised in time series applications, such as Photovoltaic (PV) power prediction, to model the stochastic volatility and reduce prediction errors. Yet the existing Wavelet Transform (WT) approach has a limitation in terms of time complexity. It requires reconstructing the decomposed components and modelling them separately and thus needs more time for reconstruction, model configuration and training. The aim of this study is to improve the efficiency of applying Wavelet Transform (WT) by proposing a new method that uses a single simplified model. Given a time series and its Wavelet Transform (WT) coefficients, it trains one model with the coefficients as features and the original time series as labels. This eliminates the need for component reconstruction and training numerous models. This work contributes to the day-ahead aggregated solar Photovoltaic (PV) power time series prediction problem by proposing and comprehensively evaluating a new approach of employing WT. The proposed approach is evaluated using 17 months of aggregated solar Photovoltaic (PV) power data from two real-world datasets. The evaluation includes the use of a variety of prediction models, including Linear Regression, Random Forest, Support Vector Regression, and Convolutional Neural Networks. The results indicate that using a coefficients-based strategy can give predictions that are comparable to those obtained using the components-based approach while requiring fewer models and less computational time.
Boosting in Machine Learning:-A Brief Overview
The post Boosting in Machine Learning:-A Brief Overview appeared first on Data Science Tutorials What do you have to lose?. Check out Data Science tutorials here Data Science Tutorials. Boosting in Machine Learning, A single predictive model, such as linear regression, logistic regression, ridge regression, etc., is the foundation of the majority of supervised machine learning methods. However, techniques such as bagging and random forests provide a wide range of models from repeated bootstrapped samples of the original dataset. The average of the predictions... Read More “Boosting in Machine Learning:-A Brief Overview” » The post Boosting in Machine Learning:-A Brief Overview appeared first on Data Science Tutorials Learn how to expert in the Data Science field with Data Science Tutorials.
7 Completely FREE R Programming Online Courses
This Free Udemy course has 3 sections. In the first section, you will learn R basics and how to download R and Rstudio. In the next section, you will learn how to code in R programming and understand functions, loops, R datasets, and R dataframes. The last section teaches how to load CSV files in R, how to apply a family of functions, how to test for normality, KNN classification, LDA(Linear Discriminant Analysis), etc. Overall, this is a good course for beginners to learn R programming basics.
SHAP: Explain Any Machine Learning Model in Python
This article is part of a series where we walk step by step in solving fintech problems with Machine Learning using "All lending club loan data". In previous articles, we prepared a dataset and built a Logistic Regression model, and we discussed the most common "ML model evaluation metrics" for a classification problem in the fintech space. This article will try to "understand" how our model decision works and what packages can help us to answer this question. Machine learning models are frequently named "black boxes". They produce highly accurate predictions.
Digital Twin and Artificial Intelligence Incorporated With Surrogate Modeling for Hybrid and Sustainable Energy Systems
Khan, Abid Hossain, Omar, Salauddin, Mushtary, Nadia, Verma, Richa, Kumar, Dinesh, Alam, Syed
Surrogate modeling has brought about a revolution in computation in the branches of science and engineering. Backed by Artificial Intelligence, a surrogate model can present highly accurate results with a significant reduction in computation time than computer simulation of actual models. Surrogate modeling techniques have found their use in numerous branches of science and engineering, energy system modeling being one of them. Since the idea of hybrid and sustainable energy systems is spreading rapidly in the modern world for the paradigm of the smart energy shift, researchers are exploring the future application of artificial intelligence-based surrogate modeling in analyzing and optimizing hybrid energy systems. One of the promising technologies for assessing applicability for the energy system is the digital twin, which can leverage surrogate modeling. This work presents a comprehensive framework/review on Artificial Intelligence-driven surrogate modeling and its applications with a focus on the digital twin framework and energy systems. The role of machine learning and artificial intelligence in constructing an effective surrogate model is explained. After that, different surrogate models developed for different sustainable energy sources are presented. Finally, digital twin surrogate models and associated uncertainties are described.
Machine Unlearning Method Based On Projection Residual
Cao, Zihao, Wang, Jianzong, Si, Shijing, Huang, Zhangcheng, Xiao, Jing
Machine learning models (mainly neural networks) are used more and more in real life. Users feed their data to the model for training. But these processes are often one-way. Once trained, the model remembers the data. Even when data is removed from the dataset, the effects of these data persist in the model. With more and more laws and regulations around the world protecting data privacy, it becomes even more important to make models forget this data completely through machine unlearning. This paper adopts the projection residual method based on Newton iteration method. The main purpose is to implement machine unlearning tasks in the context of linear regression models and neural network models. This method mainly uses the iterative weighting method to completely forget the data and its corresponding influence, and its computational cost is linear in the feature dimension of the data. This method can improve the current machine learning method. At the same time, it is independent of the size of the training set. Results were evaluated by feature injection testing (FIT). Experiments show that this method is more thorough in deleting data, which is close to model retraining.
Shuffled linear regression through graduated convex relaxation
The shuffled linear regression problem aims to recover linear relationships in datasets where the correspondence between input and output is unknown. This problem arises in a wide range of applications including survey data, in which one needs to decide whether the anonymity of the responses can be preserved while uncovering significant statistical connections. In this work, we propose a novel optimization algorithm for shuffled linear regression based on a posterior-maximizing objective function assuming Gaussian noise prior. We compare and contrast our approach with existing methods on synthetic and real data. We show that our approach performs competitively while achieving empirical running-time improvements. Furthermore, we demonstrate that our algorithm is able to utilize the side information in the form of seeds, which recently came to prominence in related problems.
Physically Meaningful Uncertainty Quantification in Probabilistic Wind Turbine Power Curve Models as a Damage Sensitive Feature
Mclean, J. H., Jones, M. R., O'Connell, B. J., Maguire, A. E, Rogers, T. J.
A wind turbines' power curve is easily accessible damage sensitive data, and as such is a key part of structural health monitoring in wind turbines. Power curve models can be constructed in a number of ways, but the authors argue that probabilistic methods carry inherent benefits in this use case, such as uncertainty quantification and allowing uncertainty propagation analysis. Many probabilistic power curve models have a key limitation in that they are not physically meaningful - they return mean and uncertainty predictions outside of what is physically possible (the maximum and minimum power outputs of the wind turbine). This paper investigates the use of two bounded Gaussian Processes in order to produce physically meaningful probabilistic power curve models. The first model investigated was a warped heteroscedastic Gaussian process, and was found to be ineffective due to specific shortcomings of the Gaussian Process in relation to the warping function. The second model - an approximated Gaussian Process with a Beta likelihood was highly successful and demonstrated that a working bounded probabilistic model results in better predictive uncertainty than a corresponding unbounded one without meaningful loss in predictive accuracy. Such a bounded model thus offers increased accuracy for performance monitoring and increased operator confidence in the model due to guaranteed physical plausibility.
A Multiple Criteria Decision Analysis based Approach to Remove Uncertainty in SMP Models
Yenduri, Gokul, Gadekallu, Thippa Reddy
Advanced AI technologies are serving humankind in a number of ways, from healthcare to manufacturing. Advanced automated machines are quite expensive, but the end output is supposed to be of the highest possible quality. Depending on the agility of requirements, these automation technologies can change dramatically. The likelihood of making changes to automation software is extremely high, so it must be updated regularly. If maintainability is not taken into account, it will have an impact on the entire system and increase maintenance costs. Many companies use different programming paradigms in developing advanced automated machines based on client requirements. Therefore, it is essential to estimate the maintainability of heterogeneous software. As a result of the lack of widespread consensus on software maintainability prediction (SPM) methodologies, individuals and businesses are left perplexed when it comes to determining the appropriate model for estimating the maintainability of software, which serves as the inspiration for this research. A structured methodology was designed, and the datasets were preprocessed and maintainability index (MI) range was also found for all the datasets expect for UIMS and QUES, the metric CHANGE is used for UIMS and QUES. To remove the uncertainty among the aforementioned techniques, a popular multiple criteria decision-making model, namely the technique for order preference by similarity to ideal solution (TOPSIS), is used in this work. TOPSIS revealed that GARF outperforms the other considered techniques in predicting the maintainability of heterogeneous automated software.
Using Knowledge Distillation to improve interpretable models in a retail banking context
Biehler, Maxime, Guermazi, Mohamed, Starck, Célim
Although the banking sector holds massive troves of data regarding its customers, products and transactions, and is no stranger to using quantitative tools to inform its decisions, two constraints usually weigh on the development of predictive models. The first one lies in the regulatory obligation to use interpretable models for a wide range of issues, with the management function being able to explain both the way a model was trained and why specific decisions have been made. Indeed, the European Banking Authority (2020) urges banking institutions to "understand the models used, and their methodology, input data, assumptions, limitations and outputs". The second has to do with the production environments available to deploy the models on. Due to the persistence of legacy systems, cost constraints or execution time limits -- think real time e-commerce fraud detection -- models may be limited to simple operations and conditions, i.e. a set of rules rather than a random forest, light computations in place of a fully fledged neural network. Modeling for retail banking use cases means dealing with both these strong customers protections -- enforced through regular audits -- and the high data volume which at times shortens the time allocated to each sample. These shackles help explain why modeling practices in retail banking departments are centered around simple and interpretable models such as the logistic regression or the (shallow) decision tree.