Information Retrieval
How Search Engines Are Killing Clever URLs
Although investors scrambled--and shelled out up to $185,000 a pop--for the chance to snatch up the new domains and profit as gatekeepers, uptake among end-users has been underwhelming. More than three years after the program's launch, roughly 26 million new generic top-level domains have been registered, compared with the 164 million registered "legacy" top-level domains. Cyrus Namazi, the vice president of domain-name services and industry engagement at ICANN, acknowledged that demand for new top-level domains won't eclipse that for legacies "any time soon." Yet Namazi believes registrations for the new extensions will continue to grow. "We are in the embryonic stages of the expansion," he said.
Omnity search engine finds documents relevant to yours โ regardless of language
With the amount of published research, patents, white papers, and other written knowledge out there, it's hard to be even reasonably sure you're aware of the goings-on around a certain topic or field. Omnity is a search engine made to make it easier by extracting the gist of documents you give it and finding related ones from a library of millions -- and now supports over a hundred languages. The process is simple and free, at least for the public-facing databases Omnity has assembled, comprising U.S. patents, SEC filings, PubMed papers, clinical trials, Library of Congress collections, and more. You upload a document or text snippet, and the system scans it, looking for the least common words and phrases -- which generally indicate things like topic, experiment type, equipment used, that sort of thing. It then looks through its own libraries to find documents with similar or related phrases that appear in a manner that suggests relevance. For example, say you put in the results of your clinical trial testing a food additive on a certain strain of mice, and found it resulted in a certain condition.
What we've learned about SEO in 2016
Since the inception of the search engine, SEO has been an important, yet often misunderstood industry. For some, these three little letters bring massive pain and frustration. For others, SEO has saved their business. One thing is for sure: having a clear and strategic search strategy is what often separates those who succeed from those who don't. As we wrap up 2016, let's take a look at how the industry has grown and shifted over the past year, and then look ahead to 2017.
Omnity's search engine uses rare word matching to find unexpected results
When it comes to search, there's Google and there's everyone else -- the company is basically synonymous with searching the internet. But Omnity, a relatively new company from San Francisco, thinks own search that's based on "semantic mapping" offers something that Google can't do. Omnity's trick is that it looks for the connections between documents on the internet based on rare words -- the theory that research that has several of the same rare words will likely be about related topics, even if that research doesn't directly link to or cite each other. Thus far, Omnity has operated primarily by selling enterprise plans to companies and educational institutions. Omnity can search not only all of the public datasets it scans (like patents, scientific, engineering and medical documents, clinical trials, case law, SEC filings and so forth) but also a company's internal documents -- for some companies, Omnity indexes 150 petabytes of data.
User Model-Based Intent-Aware Metrics for Multilingual Search Evaluation
Drutsa, Alexey, Shutovich, Andrey, Pushnyakov, Philipp, Krokhalyov, Evgeniy, Gusev, Gleb, Serdyukov, Pavel
Despite the growing importance of multilingual aspect of web search, no appropriate offline metrics to evaluate its quality are proposed so far. At the same time, personal language preferences can be regarded as intents of a query. This approach translates the multilingual search problem into a particular task of search diversification. Furthermore, the standard intent-aware approach could be adopted to build a diversified metric for multilingual search on the basis of a classical IR metric such as ERR. The intent-aware approach estimates user satisfaction under a user behavior model. We show however that the underlying user behavior models is not realistic in the multilingual case, and the produced intent-aware metric do not appropriately estimate the user satisfaction. We develop a novel approach to build intent-aware user behavior models, which overcome these limitations and convert to quality metrics that better correlate with standard online metrics of user satisfaction.
Sources of data for Search Engine
We will mainly be focusing on various sources of data that you might have to fetch or be given to build a search engine in the first place. So, if you are just an enthusiast or you have to build a professional search engine from scratch, you have come to the right place! A search engine differs from objective to objective but the core functionality remains the same โ information retrieval. Here are some of the sources of data that you might be given or you want to build a search engine for. At the heart they are all quite the same but they have quite different approaches to solving the same problem.
Tons of machine learning and data science resources that cost nothing
Tutorials, books, articles, data sets, certifications, you name it. All about data science, machine learning and related topics. You can find them with a simple keyword search: enter the keyword "free" in the DSC's search box, and here are the results. Below is a screenshot of the DSC search results page, for the keyword "free". It shows the top 6 results, out of dozens of highly relevant search results.
Spikes in search engine data predict when drugs will be recalled
Could internet searches identify dodgy drugs? A Microsoft researcher has trained an algorithm to predict whether a drug will be recalled, using queries made through Microsoft's Bing search engine. "We know that every once in a while there will be a batch of a pharmaceutical drug that will have something wrong about it," says Elad Yom-Tov at Microsoft Research in Israel. "People will start asking about that drug more often or more than they usually do." Pharmaceutical companies and regulators such as the US Food and Drug Administration (FDA) monitor drugs on the market to keep tabs on adverse effects and potential faulty batches.
ABOUT WEBSAYS - Websays
Websays is the result of 15 years of scientific investigations in Web Crawling, Automatic Learning and Text Analytics. Dr. Hugo Zaragoza, Websays' founder, is a worldwide expert in those technologies. He has worked more than 10 years as a lead researcher in Microsoft and Yahoo! in the United States, England and Spain. In 2010 Dr. Zaragoza founded Websays with the objective of applying the most cutting edge technology in information retrieval and data analytics, including various new patent pending technologies developed by Websays. Websays services focus on online reputation monitoring and social media marketing.
Microsoft researchers detect lung-cancer risks in web search logs - Next at Microsoft
Smoking cigarettes is the leading cause of lung cancer, the most common cause of cancer death in the world. But nearly 20 percent of lung-cancer diagnoses are made in people who are non-smokers. That means in addition to smoking, geographic, demographic and genetic factors play a role in the devastating disease. A project from Microsoft's research labs is exploring the feasibility of using anonymized web search data to learn more about lung-cancer risk factors and provide early warning to people who are candidates for disease screening. The findings, published Thursday in JAMA Oncology, extend research that team members published last June on the feasibility of using the text of questions people ask search engines to predict diagnoses of pancreatic cancer.