Information Retrieval
Is Google creating a voice-activated search engine for TODDLERS?
Google is potentially creating a search engine for toddlers, despite recent privacy scandals. The tech giant has filed a European patent, entitled Gamifying Voice Search Experience for Children, which gives it exclusive rights to develop the concept. Aimed at nursery-age youngsters, the prospective product would use a child-friendly bubble-interface to engage with infants. This would be separate to Google Assistant, which already allows people to conduct voice-activated searches on their devices. However, education experts have raised concerns over the risk of potential privacy violations, such as those associated with Amazon's Echo Device, plus the dangers of making children addicted to technology.
Benefits of Enabling Enterprise Search in your digital Workplace eXo
A disconnected/disengaged workforce, broken business processes and an overall decrease in efficiency represent the most recurrent challenges facing organizations today. As a result, digital workplace solutions have grown in popularity as they offer an holistic solution capable of integrating different tools and applications. A typical digital workplace includes a knowledge management system (KMS), an enterprise social network (ESN), an intranet portal, instant messaging and more. It also integrates different third party software used internally, from CRM to Human Resources Information Systems (HRIS). For better usage and efficiency, a digital workplace needs to collect data from all these data sources and make it widely accessible to users in a centralized place โ thus the importance of the enterprise search engine.
Learning to Rank Broad and Narrow Queries in E-Commerce
Devapujula, Siddhartha, Arora, Sagar, Borar, Sumit
Search is a prominent channel for discovering products on an e-commerce platform. Ranking products retrieved from search becomes crucial to address customer's need and optimize for business metrics. While learning to Rank (LETOR) models have been extensively studied and have demonstrated efficacy in the context of web search; it is a relatively new research area to be explored in the e-commerce. In this paper, we present a framework for building LETOR model for an e-commerce platform. We analyze user queries and propose a mechanism to segment queries between broad and narrow based on user's intent. We discuss different types of features - query, product and query-product and discuss challenges in using them. We show that sparsity in product features can be tackled through a denoising auto-encoder while skip-gram based word embeddings help solve the query-product sparsity issues. We also present various target metrics that can be employed for evaluating search results and compare their robustness. Further, we build and compare performances of both pointwise and pairwise LETOR models on fashion category data set. We also build and compare distinct models for broad and narrow queries, analyze feature importance across these and show that these specialized models perform better than a combined model in the fashion world.
Instagram Rolls Out In-App Local Business Profile Pages - Search Engine Journal
Instagram is introducing a new way to showcase local businesses with in-app profile pages. Raj Nijjer alerted me to this feature while providing several screenshots. As you can see in the examples below, the pages look very much like Google local knowledge panels. They have the business address, hours, contact information, and website. Of course, a link to the business's Instagram profile is featured prominently at the top of the page.
Today's customer decision journey is so complex but AI can help - Search Engine Land
Myth: "The customer journey is not as complex as it's made out to be." One thing is for sure โ the consumer decision journey is more complex than ever before. The average consumer now owns three to four devices and uses multiple online and offline channels throughout their shopping journeys. The game is changing as marketers turn to artificial intelligence, agencies and data to help them navigate new consumer behavior. Every marketer today needs to be addressing these challenges as the CDJ itself is disrupting the digital landscape.
A Novel Approach for Detection and Ranking of Trendy and Emerging Cyber Threat Events in Twitter Streams
Bose, Avishek, Behzadan, Vahid, Aguirre, Carlos, Hsu, William H.
We present a new machine learning and text information extraction approach to detection of cyber threat events in Twitter that are novel (previously non-extant) and developing (marked by significance with respect to similarity with a previously detected event). While some existing approaches to event detection measure novelty and trendiness, typically as independent criteria and occasionally as a holistic measure, this work focuses on detecting both novel and developing events using an unsupervised machine learning approach. Furthermore, our proposed approach enables the ranking of cyber threat events based on an importance score by extracting the tweet terms that are characterized as named entities, keywords, or both. We also impute influence to users in order to assign a weighted score to noun phrases in proportion to user influence and the corresponding event scores for named entities and keywords. To evaluate the performance of our proposed approach, we measure the efficiency and detection error rate for events over a specified time interval, relative to human annotator ground truth.
Real-Time Entity Resolution Made Accessible - Senzing
Knowing exactly who your customers are is an important task for security, fraud detection, marketing, and personalization. The proliferation of data sources and services has made ER very challenging in the internet age. In addition, many applications now increasingly require near real-time entity resolution.
4 chilling lessons from a tech hotline scam
While we love our smartphones, they are vulnerable to hackers. Here are some ways to keep them hacker free. Some people think they're immune to cybercriminals. "I'm not even on their radar," they think. "What are the chances that I'll get targeted? It's not like I'm famous or have zillions of dollars."
Gathering Cyber Threat Intelligence from Twitter Using Novelty Classification
Le, Ba Dung, Wang, Guanhua, Nasim, Mehwish, Babar, Ali
Preventing organizations from Cyber exploits needs timely intelligence about Cyber vulnerabilities and attacks, referred as threats. Cyber threat intelligence can be extracted from various sources including social media platforms where users publish the threat information in real time. Gathering Cyber threat intelligence from social media sites is a time consuming task for security analysts that can delay timely response to emerging Cyber threats. We propose a framework for automatically gathering Cyber threat intelligence from Twitter by using a novelty detection model. Our model learns the features of Cyber threat intelligence from the threat descriptions published in public repositories such as Common Vulnerabilities and Exposures (CVE) and classifies a new unseen tweet as either normal or anomalous to Cyber threat intelligence. We evaluate our framework using a purpose-built data set of tweets from 50 influential Cyber security related accounts over twelve months (in 2018). Our classifier achieves the F1-score of 0.643 for classifying Cyber threat tweets and outperforms several baselines including binary classification models. Our analysis of the classification results suggests that Cyber threat relevant tweets on Twitter do not often include the CVE identifier of the related threats. Hence, it would be valuable to collect these tweets and associate them with the related CVE identifier for cyber security applications.
Multi-Label Product Categorization Using Multi-Modal Fusion Models
Wirojwatanakul, Pasawee, Wangperawong, Artit
In this study, we investigated multi-modal approaches using images, descriptions, and title to categorize e-commerce products on Amazon.com. Specifically, we examined late fusion models, where the modalities are fused at the decision level. Products were each assigned multiple labels, and the hierarchy in the labels were flattened and filtered. For our individual baseline models, we modified a CNN architecture to classify the description and title, and then modified Keras' ResNet-50 to classify the images, achieving F1 scores of 77.0%, 82.7%, and 61.0%, respectively. In comparison, our tri-modal late fusion model can classify products more accurately than single modal models can, improving the F1 score to 88.2%. Each modality complemented the shortcomings of the other modalities, demonstrating that increasing the number of modalities can be an effective method for improving the accuracy of multi-label classification problems.