Information Retrieval
Record Linkage to Match Customer Names: A Probabilistic Approach
Fatemi, Bahare, Kazemi, Seyed Mehran, Poole, David
Consider the following problem: given a database of records indexed by names (e.g., name of companies, restaurants, businesses, or universities) and a new name, determine whether the new name is in the database, and if so, which record it refers to. This problem is an instance of record linkage problem and is a challenging problem because people do not consistently use the official name, but use abbreviations, synonyms, different order of terms, different spelling of terms, short form of terms, and the name can contain typos or spacing issues. We provide a probabilistic model using relational logistic regression to find the probability of each record in the database being the desired record for a given query and find the best record(s) with respect to the probabilities. Building on term-matching and translational approaches for search, our model addresses many of the aforementioned challenges and provides good results when existing baselines fail. Using the probabilities outputted by the model, we can automate the search process for a portion of queries whose desired documents get a probability higher than a trust threshold. We evaluate our model on a large real-world dataset from a telecommunications company and compare it to several state-of-the-art baselines. The obtained results show that our model is a promising probabilistic model for record linkage for names. We also test if the knowledge learned by our model on one domain can be effectively transferred to a new domain. For this purpose, we test our model on an unseen test set from the business names of the secondString dataset. Promising results show that our model can be effectively applied to unseen datasets. Finally, we study the sensitivity of our model to the statistics of datasets.
Senzing's Software for Real-Time AI for Entity Resolution to Fight Financial Crime - insideBIGDATA
Senzing, a new artificial intelligence-based (AI) software company, announced its Senzing software product to address the $14.37 billion financial fraud market. Senzing is an IBM spinout that has reinvented entity resolution, which senses who is who in real time across multiple big data sources. Senzing is disrupting the fraud solutions market by offering the first real-time, plug-and-play, AI entity resolution software product for fraud detection, insider threats and more. Now, any company can deploy Senzing to quickly and effectively detect bad actors in their big data. Senzing uses entity-centric learning and other unique techniques to pierce through falsified identities and networks to find criminals.
Metadata Enrichment of Multi-Disciplinary Digital Library: A Semantic-based Approach
Al-Natsheh, Hussein T., Martinet, Lucie, Muhlenbach, Fabrice, Rico, Fabien, Zighed, Djamel A.
In the scientific digital libraries, some papers from different research communities can be described by community-dependent keywords even if they share a semantically similar topic. Articles that are not tagged with enough keyword variations are poorly indexed in any information retrieval system which limits potentially fruitful exchanges between scientific disciplines. In this paper, we introduce a novel experimentally designed pipeline for multi-label semantic-based tagging developed for open-access metadata digital libraries. The approach starts by learning from a standard scientific categorization and a sample of topic tagged articles to find semantically relevant articles and enrich its metadata accordingly. Our proposed pipeline aims to enable researchers reaching articles from various disciplines that tend to use different terminologies. It allows retrieving semantically relevant articles given a limited known variation of search terms. In addition to achieving an accuracy that is higher than an expanded query based method using a topic synonym set extracted from a semantic network, our experiments also show a higher computational scalability versus other comparable techniques. We created a new benchmark extracted from the open-access metadata of a scientific digital library and published it along with the experiment code to allow further research in the topic.
Researchers use machine learning to search science data
As scientific datasets increase in both size and complexity, the ability to label, filter and search this deluge of information has become a laborious, time-consuming and sometimes impossible task, without the help of automated tools. With this in mind, a team of researchers from Lawrence Berkeley National Laboratory (Berkeley Lab) and UC Berkeley are developing innovative machine learning tools to pull contextual information from scientific datasets and automatically generate metadata tags for each file. Scientists can then search these files via a web-based search engine for scientific data, called Science Search, that the Berkeley team is building. As a proof-of-concept, the team is working with staff at the Department of Energy's (DOE) Molecular Foundry, located at Berkeley Lab, to demonstrate the concepts of Science Search on the images captured by the facility's instruments. A beta version of the platform has been made available to Foundry researchers.
Berkeley Lab researchers use machine learning to search science data
IMAGE: This is a screenshot of the Science Search interface. In this case, the user did an image search of nanoparticles. As scientific datasets increase in both size and complexity, the ability to label, filter and search this deluge of information has become a laborious, time-consuming and sometimes impossible task, without the help of automated tools. With this in mind, a team of researchers from Lawrence Berkeley National Laboratory (Berkeley Lab) and UC Berkeley are developing innovative machine learning tools to pull contextual information from scientific datasets and automatically generate metadata tags for each file. Scientists can then search these files via a web-based search engine for scientific data, called Science Search, that the Berkeley team is building.
Researchers Use Machine Learning to Search Science Data
In this case, the user performed an image search for nanoparticles. As scientific datasets increase in both size and complexity, the ability to label, filter and search this deluge of information has become a laborious, time-consuming and sometimes impossible task, without the help of automated tools. With this in mind, a team of researchers from the Department of Energy's Lawrence Berkeley National Laboratory (Berkeley Lab) and UC Berkeley are developing innovative machine learning tools to pull contextual information from scientific datasets and automatically generate metadata tags for each file. Scientists can then search these files via a web-based search engine for scientific data, called Science Search, that the Berkeley team is building. As a proof-of-concept, the team is working with staff at Berkeley Lab's Molecular Foundry, to demonstrate the concepts of Science Search on the images captured by the facility's instruments.
How To Create Natural Language Semantic Search For Arbitrary Objects With Deep Learning
The power of modern search engines is undeniable: you can summon knowledge from the internet at a moment's notice. There are many situations where search is relegated to strict keyword search, or when the objects aren't text, search may not be available. Furthermore, strict keyword search doesn't allow the user to search semantically, which means information is not as discoverable. Today, we share a reproducible, minimally viable product that illustrates how you can enable semantic search for arbitrary objects! Concretely, we will show you how to create a system that searches python code semantically -- but this approach can be generalized to other entities (such as pictures or sound clips).
Google Has Removed Over 80% of Hacked Sites from Search Results - Search Engine Journal
Google has released new details about about its spam fighting efforts, revealing that more than 80% of hacked sites have been detected and removed from search results. The search giant plans to continue its efforts by working directly with popular content management systems to fight back against those who compromise forums and comment sections with spam. "Last year, we focused a great deal of effort on reducing the impact on users from hacked websites, and were able to detect and remove more than 80 percent of compromised sites from search results. We're also working closely with many providers of popular content management systems like WordPress and Joomla to help them fight spammers that abuse forums and comment sections." Here are some other notable stats from Google's recent announcement.
TrQuery: An Embedding-based Framework for Recommanding SPARQL Queries
Zhang, Lijing, Zhang, Xiaowang, Feng, Zhiyong
In this paper, we present an embedding-based framework (TrQuery) for recommending solutions of a SPARQL query, including approximate solutions when exact querying solutions are not available due to incompleteness or inconsistencies of real-world RDF data. Within this framework, embedding is applied to score solutions together with edit distance so that we could obtain more fine-grained recommendations than those recommendations via edit distance. For instance, graphs of two querying solutions with a similar structure can be distinguished in our proposed framework while the edit distance depending on structural difference becomes unable. To this end, we propose a novel score model built on vector space generated in embedding system to compute the similarity between an approximate subgraph matching and a whole graph matching. Finally, we evaluate our approach on large RDF datasets DBpedia and YAGO, and experimental results show that TrQuery exhibits an excellent behavior in terms of both effectiveness and efficiency.
Supercharging your SEO with AI: Insights, automation and personalization - Search Engine Land
Recently, I had the pleasure of presenting at SMX London on Supercharging your SEO with AI and thought I would share some of the insights with Search Engine Land readers. Google made global headlines with the demonstration of its new Duplex at this year's I/O developers conference. This artificial intelligence (AI) system can "converse" in natural language with people to schedule an appointment at a hair salon or book a table at a restaurant, for example. To pass the Turing Test, AI must behave in a manner indistinguishable from that of a human. To many, Google Duplex has proven that it can pass this test, but in truth, we are only seeing the beginnings of its future potential.