Information Retrieval
Using Search Queries to Understand Health Information Needs in Africa
Abebe, Rediet, Hill, Shawndra, Vaughan, Jennifer Wortman, Small, Peter M., Schwartz, H. Andrew
The lack of comprehensive, high-quality health data in developing nations creates a roadblock for combating the impacts of disease. One key challenge is understanding the health information needs of people in these nations. Without understanding people's everyday needs, concerns, and misconceptions, health organizations and policymakers lack the ability to effectively target education and programming efforts. In this paper, we propose a bottom-up approach that uses search data from individuals to uncover and gain insight into health information needs in Africa. We analyze Bing searches related to HIV/AIDS, malaria, and tuberculosis from all 54 African nations. For each disease, we automatically derive a set of common search themes or topics, revealing a wide-spread interest in various types of information, including disease symptoms, drugs, concerns about breastfeeding, as well as stigma, beliefs in natural cures, and other topics that may be hard to uncover through traditional surveys. We expose the different patterns that emerge in health information needs by demographic groups (age and sex) and country. We also uncover discrepancies in the quality of content returned by search engines to users by topic. Combined, our results suggest that search data can help illuminate health information needs in Africa and inform discussions on health policy and targeted education efforts both on- and offline.
Consistent Position Bias Estimation without Online Interventions for Learning-to-Rank
Agarwal, Aman, Zaitsev, Ivan, Joachims, Thorsten
Presentation bias is one of the key challenges when learning from implicit feedback in search engines, as it confounds the relevance signal with uninformative signals due to position in the ranking, saliency, and other presentation factors. While it was recently shown how counterfactual learning-to-rank (LTR) approaches \cite{Joachims/etal/17a} can provably overcome presentation bias if observation propensities are known, it remains to show how to accurately estimate these propensities. In this paper, we propose the first method for producing consistent propensity estimates without manual relevance judgments, disruptive interventions, or restrictive relevance modeling assumptions. We merely require that we have implicit feedback data from multiple different ranking functions. Furthermore, we argue that our estimation technique applies to an extended class of Contextual Position-Based Propensity Models, where propensities not only depend on position but also on observable features of the query and document. Initial simulation studies confirm that the approach is scalable, accurate, and robust.
This Vietnamese Browser & Search Engine Is Daring Google To Step-Up Its Game
Cแปc Cแปc's browser has gained significant market share in Vietnam This was the response I received when I started chatting with the only other customer at a restaurant in Hanoi, Vietnam exactly four days into my journey traveling and meeting with startups around the world. I had been contemplating how best to break into the Vietnamese startup ecosystem, and this chance meeting proved to be the answer. It turned out the tall, friendly Russian I was speaking with was Victor Lavrenko, CEO of Cแปc Cแปc, Vietnam's leading local browser and search engine! Over the next few days, I got to visit Cแปc Cแปc's office and learn more from Lavrenko and CMO Kristina Melentieva about the company and why it's worth keeping an eye on. Only once I left the Cแปc Cแปc office--with its ping pong table, vibrant color scheme, and full-wall ideation whiteboard--and ventured back onto the street did I fully appreciate the obvious: this company is not operating (and flourishing) in any number of internationally recognized technology startup hubs, but in the chaotic and volatile center of Hanoi.
Real Estate Search Engine Powered by Artificial Intelligence
ITRealty intelligently analyzes the multiple listing services (MLS) that brokers and agents use to find/list properties, establish contractual offers of compensation among brokers, and accumulate and disseminate information to enable appraisals. "It does not matter anymore how far in the past comparable properties were sold", says Yuriy Setko, "Our algorithms will analyze where the market was for that particular type of property, in that particular neighborhood in the past, and apply the time adjustment percentage to the selling price to give you that property's market price as if it was sold yesterday." You usually have a good number of comparables to see if the asking price is right". It also helps to precisely determine the market price when putting up a property for sale. ITRealty tracks price drops on MLS, along with other listing analysis algorithms, to find "motivated sellers", as well as drawing supplementary data on real estate not found on MLS.
What do Google and a toddler have in common? Both need to learn good listening skills. - Search Engine Land
At the Sixth International Conference on Learning Representations, Jannis Bulian and Neil Houlsby, researchers at Google AI, presented a paper that shed light on new methods they're testing to improve search results. While publishing a paper certainly doesn't mean the methods are being used, or even will be, it likely increases the odds when the results are highly successful. And when those methods also combine with other actions Google is taking, one can be almost certain. I believe this is happening, and the changes are significant for search engine optimization specialists (SEOs) and content creators. Let's start with the basics and look topically at what's being discussed.
How San Quentin Inmates Built JOLT, a Search Engine for Prison
Marcellino Ornelas had been in and out of juvenile hall seven times by the time he finally went to prison at the age of 19 for assault with a firearm. He'd already been kicked out of high school and was working, he says, as the "local drug dealer," with a side gig at a Ross department store. In the past, every time he got out, he'd start dealing soon after. "It was like, this is how I make money. This is who my friends are," Ornelas says.
A Systematic Classification of Knowledge, Reasoning, and Context within the ARC Dataset
Boratko, Michael, Padigela, Harshit, Mikkilineni, Divyendra, Yuvraj, Pritish, Das, Rajarshi, McCallum, Andrew, Chang, Maria, Fokoue-Nkoutche, Achille, Kapanipathi, Pavan, Mattei, Nicholas, Musa, Ryan, Talamadupula, Kartik, Witbrock, Michael
The recent work of Clark et al. (2018) introduces the AI2 Reasoning Challenge (ARC) and the associated ARC dataset that partitions open domain, complex science questions into an Easy Set and a Challenge Set. That paper includes an analysis of 100 questions with respect to the types of knowledge and reasoning required to answer them; however, it does not include clear definitions of these types, nor does it offer information about the quality of the labels. We propose a comprehensive set of definitions of knowledge and reasoning types necessary for answering the questions in the ARC dataset. Using ten annotators and a sophisticated annotation interface, we analyze the distribution of labels across the Challenge Set and statistics related to them. Additionally, we demonstrate that although naive information retrieval methods return sentences that are irrelevant to answering the query, sufficient supporting text is often present in the (ARC) corpus. Evaluating with human-selected relevant sentences improves the performance of a neural machine comprehension model by 42 points.
A Hidden Instagram Feature Shows Users Time Spent in the App - Search Engine Journal
Instagram is testing a hidden "Usage Insights" feature, which shows users how much time they spend in the app. This feature was discovered by a computer science student who has a history of uncovering new features in Instagram before they're rolled out to the public. All that we have to go on at this time is the screenshot shared by Jade M. Wong, so it's unclear how detailed these insights are. Instagram is testing "Usage Insights" to show the amount of time users have spent on the app Be self-aware or be prepared to be ashamed for Instagram addiction pic.twitter.com/WzyRGWIOgZ It's also unclear whether or not this feature will be widely released, although it's not something that's out of the realm of possibility. Earlier this week, Google announced it will be rolling out a feature designed to keep users informed about how much time they're spending on YouTube.
What negative SEO is and is not - Search Engine Land
Today we are starting a six-part series on Negative SEO. The series will be broken into three areas and will show how negative search engine optimization (SEO) has an effect on links, content and user signals. Positive SEO under this broader view would be any tactic performed with the intent to positively impact rankings for a uniform resource locator (URL), and possibly its host domain, by manipulating a variable within the links, content or user signals areas. Negative SEO would be any tactic performed with the intent to negatively impact rankings for a URL, and possibly its host domain, by manipulating a variable within the links, content or user signal buckets. If you can accidentally hurt your rankings by shifting a variable, then it would logically suggest that an external entity shifting that same variable associated with your site could result in a ranking decrease or outright deindexation.
Exploiting Textual and Citation Information to Identify and Summarize Influential Publications
Zahran, Mohamed A. (Purdue University) | Ebaid, Amr (Purdue University)
Given a group of publications, we investigate the prob- lem of identifying the papers with the most impact on others. We refer to these papers as influential in the sense that they introduce new concepts and language that will affect how future articles are written. In this pa- per we propose weighted PageRank algorithm that uses textual information from articles and information from citation graph to rank the impact of publications, then we automatically summarize these publications and ex- tract important keywords. We show that using our algo- rithm outperforms default citation-based techniques in ranking influential papers (those which won best paper award) with no less than 2% in F1-score and NDCG. We also show that our algorithm outperforms previous graph-based keyword extraction techniques with no less than 1.5% in F1-score.