Goto

Collaborating Authors

 data science approach


Council Post: Top Five Data Science Trends That Made An Impact In 2022

#artificialintelligence

With the increasing amount of data and the increasing awareness of data-driven culture, global businesses strive to adopt a data science approach. Undoubtedly, data-driven intelligence has become the highest parameter to succeed in the digital world. However, Covid changed the world overnight. Most data science models became useless--at least for some time. Everyone raced to retrain and redeploy their existing data science models.


Accelerating drug discovery from bed to benchside - Healthskouts

#artificialintelligence

Silicon Valley giant NVIDIA is teaming up with pharma company AstraZeneca and the University of Florida on new artificial intelligence research projects aimed at boosting drug discovery and patient care. April 21, NVIDIA and AstraZeneca revealed a new drug-discovery model called MegaMoIBART, which is aimed at "reaction prediction, molecular optimization and de novo molecular generation." MegaMoIBART will be deployable on NVIDIA's platform for computational drug discovery, known as Clara Discovery, and will use a new kind of technology called transformer neural networks. This is the new breed of press releases flooding the domain of drug discovery, until recently the field of pure pharma & life sciences companies, medical chemistry procedures and very time-consuming biologic research. I used to discover and develop novel candidate drugs myself.


Data Science Approach from Scratch: An Easy Explanation - CouponED

#artificialintelligence

The key here is having the best understanding of the problem. Welcome to the ultimate course on Data Science Approach from Scratch!!! This course is your Best Resource for learning the use of Data Science.


Identifying Semantically Duplicate Questions Using Data Science Approach: A Quora Case Study

arXiv.org Machine Learning

Identifying semantically identical questions on, Question and Answering social media platforms like Quora is exceptionally significant to ensure that the quality and the quantity of content are presented to users, based on the intent of the question and thus enriching overall user experience. Detecting duplicate questions is a challenging problem because natural language is very expressive, and a unique intent can be conveyed using different words, phrases, and sentence structuring. Machine learning and deep learning methods are known to have accomplished superior results over traditional natural language processing techniques in identifying similar texts. In this paper, taking Quora for our case study, we explored and applied different machine learning and deep learning techniques on the task of identifying duplicate questions on Quora's dataset. By using feature engineering, feature importance techniques, and experimenting with seven selected machine learning classifiers, we demonstrated that our models outperformed previous studies on this task. Xgboost model with character level term frequency and inverse term frequency is our best machine learning model that has also outperformed a few of the Deep learning baseline models. We applied deep learning techniques to model four different deep neural networks of multiple layers consisting of Glove embeddings, Long Short Term Memory, Convolution, Max pooling, Dense, Batch Normalization, Activation functions, and model merge. Our deep learning models achieved better accuracy than machine learning models. Three out of four proposed architectures outperformed the accuracy from previous machine learning and deep learning research work, two out of four models outperformed accuracy from previous deep learning study on Quora's question pair dataset, and our best model achieved accuracy of 85.82% which is close to Quora state of the art accuracy.


A Data Science Approach for Honeypot Detection in Ethereum

arXiv.org Machine Learning

Ethereum smart contracts have recently drawn a considerable amount of attention from the media, the financial industry and academia. With the increase in popularity, malicious users found new opportunities to profit from deceiving newcomers. Consequently, attackers started luring other attackers into contracts that seem to have exploitable flaws, but that actually contain a complex hidden trap that in the end benefits the contract creator. This kind of contracts are known in the blockchain community as Honeypots. A recent study, proposed to investigate this phenomenon by focusing on the contract bytecode using symbolic analysis. In this paper, we present a data science approach based on the contract transaction behavior. We create a partition of all the possible cases of fund movement between the contract creator, the contract, the sender of the transaction and other participants. We calculate the frequency of every case per contract, and extract as well other contract features and transaction aggregated features. We use the collected information to train machine learning models that classify contracts as honeypot or non-honeypots, and also measure how well they perform when classifying unseen honeypot types. We compare our results with the bytecode analysis method using labels from a previous study, and discuss in which cases each solution has advantages over the other.


"Data Science Approach to Reducing Employee Attrition" from #JobsOfFuture Discussion future of work, worker and workplace! by TAO.ai on Apple Podcasts

#artificialintelligence

We are unable to find iTunes on your computer. To download and subscribe to #JobsOfFuture Discussion future of work, worker and workplace! Click I Have iTunes to open it now. To listen to an audio podcast, mouse over the title and click Play. Open iTunes to download and subscribe to podcasts.


"Data Science Approach to Reducing Employee Attrition" from #JobsOfFuture Discussion future of work, worker and workplace! by TAO.ai on Apple Podcasts

#artificialintelligence

We are unable to find iTunes on your computer. To download and subscribe to #JobsOfFuture Discussion future of work, worker and workplace! Click I Have iTunes to open it now. To listen to an audio podcast, mouse over the title and click Play. Open iTunes to download and subscribe to podcasts.


AI: It's All About the Data: The Shift from Computer Science to Data Science – Part 2

#artificialintelligence

The newest data science approach to managing and optimizing virtual infrastructures applies the AI discipline of machine learning (ML). Rather than monitoring individual components in the traditional computer science way, ML tools analyze the behavior of interrelated components. They track the normal patterns of these complex behaviors as they change over time. Machine learning-based analytics tools automatically identify the root causes of performance issues and recommend the steps needed to fix them. This shift to a data-centric, behavior-based approach has major implications that significantly empower IT professionals.


A Data Science Approach for Device Level Operational State Classification Using Real Time Energy Data

#artificialintelligence

Recent developments in energy management systems and the IoT (Internet of Things), have enabled easy, and low cost visibility of real time energy consumption data of not only main power lines but also individual devices. For anyone skilled in the art of energy management, it is obvious that such data contains incredible value that can help facility managers significantly increase the operational and energy efficiency of their sites. However, due to the shortage and cost of analytical resources, it is always a great challenge to practically and easily deliver such valuable insights out of so much data. As more and more devices are being monitored, the task becomes nearly impossible to manage manually. An article which I recently published as part of the latest research work we're doing in Panoramic Power, introduces an innovative data-science approach that helps automatically generate actionable energy and operational efficiency insights out of real time device level energy consumption data, using machine learning techniques.


The Green Jacket Goes To…? Using a Data Science Approach to Picking the Winner of The Masters

#artificialintelligence

Spring is the time when many sports fans are glued to their brackets in hopes of asserting their ability to correctly select amongst the 150 quintillion permutations of Teams who will win the NCAA Basketball Championship (see our blog post on this subject). My personal highlight of the spring sports calendar is the Masters Golf Tournament, which is held every year at the Augusta National Golf Club in Georgia. The professional golf schedule contains four major tournaments each year: The Masters, The US Open, The Open Championship, and The PGA Championship. Of these tournaments, only the Masters is played on the same course every year and its champion is awarded the iconic Masters Champion green jacket. Having participated in a number of fantasy sports leagues and being a Data Scientist at MapR gives me a unique perspective on my approach to choosing who I think will most likely "win" the tournament.