Education
At the crossroads of language, technology, and empathy
Rujul Gandhi's love of reading blossomed into a love of language at age 6, when she discovered a book at a garage sale called "What's Behind the Word?" With forays into history, etymology, and language genealogies, the book captivated Gandhi, who as an MIT senior remains fascinated with words and how we use them. Growing up partially in the U.S. and mostly in India, Gandhi was surrounded by a variety of languages and dialects. When she moved to India at age 8, she could already see how knowing the Marathi language allowed her to connect more easily to her classmates -- an early lesson in how language shapes our human experiences. Initially thinking she might want to study creative writing or theater, Gandhi first learned about linguistics as its own field of study through an online course in ninth grade.
A Little About Me -- Amena Khatun
My name is Amena Khatun, and I currently live in Australia with my partner and our son. I am working as a'Postdoctoral Fellow' at Queensland University and Technology (QUT), Brisbane, Australia. My research interest is computer vision, deep learning, person re-identification, and security surveillance. In February 2017, my Ph.D. journey started in computer vision and deep learning at QUT. I am so thankful for the Australian Government RTP Scholarship, QUT HDR tuition Fees Sponsorship, and QUT Top-up Scholarship.
Learning with Subset Stacking
Birbil, S. Ilker, Yildirim, Sinan, Gokalp, Kaya, Akyuz, Hakan
We propose a new algorithm that learns from a set of input-output pairs. Our algorithm is designed for populations where the relation between the input variables and the output variable exhibits a heterogeneous behavior across the predictor space. The algorithm starts with generating subsets that are concentrated around random points in the input space. This is followed by training a local predictor for each subset. Those predictors are then combined in a novel way to yield an overall predictor. We call this algorithm "LEarning with Subset Stacking" or LESS, due to its resemblance to method of stacking regressors. We compare the testing performance of LESS with the state-of-the-art methods on several datasets. Our comparison shows that LESS is a competitive supervised learning method. Moreover, we observe that LESS is also efficient in terms of computation time and it allows a straightforward parallel implementation.
Up to 100x Faster Data-free Knowledge Distillation
Fang, Gongfan, Mo, Kanya, Wang, Xinchao, Song, Jie, Bei, Shitao, Zhang, Haofei, Song, Mingli
Data-free knowledge distillation (DFKD) has recently been attracting increasing attention from research communities, attributed to its capability to compress a model only using synthetic data. Despite the encouraging results achieved, state-of-the-art DFKD methods still suffer from the inefficiency of data synthesis, making the data-free training process extremely time-consuming and thus inapplicable for large-scale tasks. In this work, we introduce an efficacious scheme, termed as FastDFKD, that allows us to accelerate DFKD by a factor of orders of magnitude. At the heart of our approach is a novel strategy to reuse the shared common features in training data so as to synthesize different data instances. Unlike prior methods that optimize a set of data independently, we propose to learn a meta-synthesizer that seeks common features as the initialization for the fast data synthesis. As a result, FastDFKD achieves data synthesis within only a few steps, significantly enhancing the efficiency of data-free training. Experiments over CIFAR, NYUv2, and ImageNet demonstrate that the proposed FastDFKD achieves 10$\times$ and even 100$\times$ acceleration while preserving performances on par with state of the art.
DeepFIB: Self-Imputation for Time Series Anomaly Detection
Liu, Minhao, Xu, Zhijian, Xu, Qiang
Time series (TS) anomaly detection (AD) plays an essential role in various applications, e.g., fraud detection in finance and healthcare monitoring. Due to the inherently unpredictable and highly varied nature of anomalies and the lack of anomaly labels in historical data, the AD problem is typically formulated as an unsupervised learning problem. The performance of existing solutions is often not satisfactory, especially in data-scarce scenarios. To tackle this problem, we propose a novel self-supervised learning technique for AD in time series, namely \emph{DeepFIB}. We model the problem as a \emph{Fill In the Blank} game by masking some elements in the TS and imputing them with the rest. Considering the two common anomaly shapes (point- or sequence-outliers) in TS data, we implement two masking strategies with many self-generated training samples. The corresponding self-imputation networks can extract more robust temporal relations than existing AD solutions and effectively facilitate identifying the two types of anomalies. For continuous outliers, we also propose an anomaly localization algorithm that dramatically reduces AD errors. Experiments on various real-world TS datasets demonstrate that DeepFIB outperforms state-of-the-art methods by a large margin, achieving up to $65.2\%$ relative improvement in F1-score.
Artificial Intelligence in the Biopharma Lifecycle
This program will provide U.S. and European perspectives on intellectual property considerations for use of artificial intelligence ("AI") in the lifecycle of biopharmaceutical products. Specific topics will include AI networks and tools, the AI-biopharma space, patent portfolio considerations, and IP transactional considerations. We hope you will be able to participate in this timely webinar. This program has been approved for 1.00 hour of general distance learning credit by the Pennsylvania CLE Board. This program has also been approved for 1.00 hour of general credit by the State Bar of California and 1.00 hour of areas of professional practice credit (including transitional) by the New York State CLE Board.
Data Science
Interviews are the most challenging part of getting any job especially for Data Scientist and Machine Learning Engineer roles where you are tested on Machine Learning and Deep Learning concepts. So, Given below is a short quiz that consists of 25 Questions consisting of MCQs(One or more correct), True-False, and Integer Type Questions to check your knowledge. The ideal time devoted to this quiz should be around 50 min. What does this course offer you? Every question is associated with a knowledge area based on Data Scientist and Machine Learning.
Machine Learning Practical: 6 Real-World Applications
So you know the theory of Machine Learning and know how to create your first algorithms. There are tons of courses out there about the underlying theory of Machine Learning which don't go any deeper – into the applications. This course is not one of them. Are you ready to apply all of the theory and knowledge to real life Machine Learning challenges? We gathered best industry professionals with tons of completed projects behind.
10-best-machine-learning-start-ups-to-watch-in-2022
This is the list of the 10 most exciting machine learning start-ups you should be following in 2022. Artificial Intelligence has been a hot area of innovation in recent years and ML is one of the major sections of the whole AI arena. ML refers to the development of intelligent algorithms and statistical modeling that allow for further programming improvement without having to code them explicitly. Machine learning can make a predictive analysis app more precise over time, for instance. ML is not without its problems.