Information Retrieval
Keyword Extraction API - BytesView
Keyword extraction also known as keyword detection is a machine learning technique that can help you automate the identification and extraction of relevant information from unstructured text data. BytesView's efficient keyword extraction tool can analyze unstructured text including customer feedback, emails, surveys, social media posts, etc. Pre-define tags to identify topical content, business intelligence, customer opinions, and recurring tickets.
Google Continues To Pay Apple Billions To Remain Safari's Default Search Engine
According to a report from Ped30, they have gotten their hands on an investor's note from Bernstein's analysts where they are claiming that Google is now paying Apple as much as $15 billion in 2021 to remain Safari's default search. This is higher than what Google had paid Apple in 2020 at $10 billion, and it seems that this figure is only expected to grow. According to the analysts, "We now estimate that Google's payments to AAPL to be the default search engine on iOS were $10B in FY 20, higher than our prior published model estimate of $8B. Recent disclosures in Apple's public filings as well as a bottom-up analysis of Google's TAC (traffic acquisition costs) payments each point us to this figureโฆWe now forecast that Google's payments to Apple might be nearly $15B in FY 21, contribute an amazing 850 bps to Services growth YoY, and amount to 9% of company gross profits." They go on to estimate that this figure will jump to $18-$20 billion in 2022, and the reason behind the increase in payments is because Google wants to ensure that Microsoft (and other competitors) don't outbid them.
sigmoidF1: A Smooth F1 Score Surrogate Loss for Multilabel Classification
Bรฉnรฉdict, Gabriel, Koops, Vincent, Odijk, Daan, de Rijke, Maarten
Multiclass multilabel classification refers to the task of attributing multiple labels to examples via predictions. Current models formulate a reduction of that multilabel setting into either multiple binary classifications or multiclass classification, allowing for the use of existing loss functions (sigmoid, cross-entropy, logistic, etc.). Empirically, these methods have been reported to achieve good performance on different metrics (F1 score, Recall, Precision, etc.). Theoretically though, the multilabel classification reductions does not accommodate for the prediction of varying numbers of labels per example and the underlying losses are distant estimates of the performance metrics. We propose a loss function, sigmoidF1. It is an approximation of the F1 score that (I) is smooth and tractable for stochastic gradient descent, (II) naturally approximates a multilabel metric, (III) estimates label propensities and label counts. More generally, we show that any confusion matrix metric can be formulated with a smooth surrogate. We evaluate the proposed loss function on different text and image datasets, and with a variety of metrics, to account for the complexity of multilabel classification evaluation. In our experiments, we embed the sigmoidF1 loss in a classification head that is attached to state-of-the-art efficient pretrained neural networks MobileNetV2 and DistilBERT. Our experiments show that sigmoidF1 outperforms other loss functions on four datasets and several metrics. These results show the effectiveness of using inference-time metrics as loss function at training time in general and their potential on non-trivial classification problems like multilabel classification.
QUEACO: Borrowing Treasures from Weakly-labeled Behavior Data for Query Attribute Value Extraction
Zhang, Danqing, Li, Zheng, Cao, Tianyu, Luo, Chen, Wu, Tony, Lu, Hanqing, Song, Yiwei, Yin, Bing, Zhao, Tuo, Yang, Qiang
We study the problem of query attribute value extraction, which aims to identify named entities from user queries as diverse surface form attribute values and afterward transform them into formally canonical forms. Such a problem consists of two phases: {named entity recognition (NER)} and {attribute value normalization (AVN)}. However, existing works only focus on the NER phase but neglect equally important AVN. To bridge this gap, this paper proposes a unified query attribute value extraction system in e-commerce search named QUEACO, which involves both two phases. Moreover, by leveraging large-scale weakly-labeled behavior data, we further improve the extraction performance with less supervision cost. Specifically, for the NER phase, QUEACO adopts a novel teacher-student network, where a teacher network that is trained on the strongly-labeled data generates pseudo-labels to refine the weakly-labeled data for training a student network. Meanwhile, the teacher network can be dynamically adapted by the feedback of the student's performance on strongly-labeled data to maximally denoise the noisy supervisions from the weak labels. For the AVN phase, we also leverage the weakly-labeled query-to-attribute behavior data to normalize surface form attribute values from queries into canonical forms from products. Extensive experiments on a real-world large-scale E-commerce dataset demonstrate the effectiveness of QUEACO.
Azure Synapse Analytics Serverless SQL Pool Guidelines
With the introduction of the serverless SQL pool as a part of Azure Synapse Analytics, Microsoft has provided a very cost-efficient and convenient way to drive value from data residing in lakes using simple T-SQL statements. It enables you to easily build logical analytical models by querying and joining data across heterogeneous sources making the development of complex data integration pipelines obsolete in many cases. To use it, you don't even need to explicitly provision it beforehand due to its serverless nature, it is per default part of an Azure Synapse Analytics workspace. All you have to do is query data in an on-demand fashion in which you get charged according to the amount of data your queries need to process. Yet, the flexibility provided in terms of how data can be stored and queried require you to stick to some conventions for properly applying all its features and functionalities. Otherwise, the once promising serverless query engine can end up causing lots of costs together with a poor performance.
Towards Personalized and Human-in-the-Loop Document Summarization
The ubiquitous availability of computing devices and the widespread use of the internet have generated a large amount of data continuously. Therefore, the amount of available information on any given topic is far beyond humans' processing capacity to properly process, causing what is known as information overload. To efficiently cope with large amounts of information and generate content with significant value to users, we require identifying, merging and summarising information. Data summaries can help gather related information and collect it into a shorter format that enables answering complicated questions, gaining new insight and discovering conceptual boundaries. This thesis focuses on three main challenges to alleviate information overload using novel summarisation techniques. It further intends to facilitate the analysis of documents to support personalised information extraction. This thesis separates the research issues into four areas, covering (i) feature engineering in document summarisation, (ii) traditional static and inflexible summaries, (iii) traditional generic summarisation approaches, and (iv) the need for reference summaries. We propose novel approaches to tackle these challenges, by: i)enabling automatic intelligent feature engineering, ii) enabling flexible and interactive summarisation, iii) utilising intelligent and personalised summarisation approaches. The experimental results prove the efficiency of the proposed approaches compared to other state-of-the-art models. We further propose solutions to the information overload problem in different domains through summarisation, covering network traffic data, health data and business process data.
How Search Engines Use Machine Learning: 9 Things We Know For Sure
Tech giants are investing heavily in machine learning. In 2019, Microsoft invested in 11 artificial intelligence (AI) startups, with $1 billion for OpenAI alone. In that same year, Intel Capital made 19 investments, and Google Ventures made 16 investments. That huge influx of capital means that AI computing power is making rapid advancements in a range of sectors from healthcare to construction to marketing and search engine optimization. However, before we get into the implications of machine learning for SEO professionals, let's define what we mean by AI.
Web image search engine based on LSH index and CNN Resnet50
Parola, Marco, Nannini, Alice, Poleggi, Stefano
To implement a good Content Based Image Retrieval (CBIR) system, it is essential to adopt efficient search methods. One way to achieve this results is by exploiting approximate search techniques. In fact, when we deal with very large collections of data, using an exact search method makes the system very slow. In this project, we adopt the Locality Sensitive Hashing (LSH) index to implement a CBIR system that allows us to perform fast similarity search on deep features. Specifically, we exploit transfer learning techniques to extract deep features from images; this phase is done using two famous Convolutional Neural Networks (CNNs) as features extractors: Resnet50 and Resnet50v2, both pre-trained on ImageNet. Then we try out several fully connected deep neural networks, built on top of both of the previously mentioned CNNs in order to fine-tuned them on our dataset. In both of previous cases, we index the features within our LSH index implementation and within a sequential scan, to better understand how much the introduction of the index affects the results. Finally, we carry out a performance analysis: we evaluate the relevance of the result set, computing the mAP (mean Average Precision) value obtained during the different experiments with respect to the number of done comparison and varying the hyper-parameter values of the LSH index.
AI DBT Impact on Mammography Post-breast Therapy
AI-CAD marked axillary lymph node and region in right upper outer quadrant (arrows and thin line outlining both sites) and assigned an abnormality score of 28%. August 11, 2021 -- According to an open-access Editor's Choice article in the American Journal of Roentgenology (AJR), artificial intelligence-based computer-aided detection (AI-CAD) can be a practical addition for lowering false-positive findings when performing post-breast conserving therapy (BCT) surveillance mammography. "After BCT, adjunct digital breast tomosynthesis (DBT) or AI-CAD reduced recall rates and improved accuracy in the ipsilateral and contralateral breasts compared with digital mammography (DM)," wrote lead investigator Jung Hyun Yoon. "In the ipsilateral breast, addition of AI-CAD resulted in lower recall rate and higher accuracy than addition of DBT." Yoon and colleagues' single-center retrospective study included 314 women (mean age, 53.2 years; 4 with bilateral breast cancer) who underwent BCT followed by DBT (mean interval from surgery to DBT, 15.2 months). Three breast radiologists independently reviewed images in three sessions: DM, DM with DBT, and DM with AI-CAD.