Information Retrieval
Ambiguity Aware Arabic Document Indexing and Query Expansion: A Morphological Knowledge Learning-Based Approach
Soudani, Nadia (La Manouba University) | Bounhas, Ibrahim (La Manouba University) | Babis, Sawssen Ben (La Manouba University)
In this paper, we propose a morphology-based Arabic Information Retrieval (IR) system. Arabic is an inflectional and derivational language and Arabic texts are highly ambiguous at the morphological level. However, short diacritics have a central role in understanding Arabic texts. That is, we propose to build a morphological knowledge base from huge vocalized corpora to reduce the ambiguity of Arabic documents. This base may be used both for the morphological indexing of queries and documents and to the morphological enrichment of queries. Indeed, it stores (i) the morpho-syntactic attributes of Arabic words; and, (ii) the morphological relations between Arabic tokens. It also represents the Arabic lexicon at several levels (e.g. stems, lemmas and words). We focus on morphological analysis and disambiguation and its impact in information retrieval. We perform experiments, which try to study the problem of indexing units and morphology-based query expansion in Arabic IR.
Search Engine Optimization Tutorial for Beginners
This course centers around the technical steps you need to take to put your online assets (website, blog site, online store, etc.) in the best possible light in the eyes of search engines, more specifically Google and Bing-Yahoo. This course covers thing you absolutely must do to have even a shot at getting to page one of these search engines organically. I follow these lectures with additional guidance to help you construct and display your web page assets so that they not only pass scrutiny when being crawled by "Spiders" in the service of these search engines, but that they receive high-marks from these activities - which will hep you to place higher in the search engine rankings when organic searches are conducted by the public looking for online information. There are many "website designers" out there today completing sites using templates, widgets, etc. And the whole world is out there using similar keywords and keyword phrases trying to get found.
Cross-lingual Document Retrieval using Regularized Wasserstein Distance
Balikas, Georgios, Laclau, Charlotte, Redko, Ievgen, Amini, Massih-Reza
Many information retrieval algorithms rely on the notion of a good distance that allows to efficiently compare objects of different nature. Recently, a new promising metric called Word Mover's Distance was proposed to measure the divergence between text passages. In this paper, we demonstrate that this metric can be extended to incorporate term-weighting schemes and provide more accurate and computationally efficient matching between documents using entropic regularization. We evaluate the benefits of both extensions in the task of cross-lingual document retrieval (CLDR). Our experimental results on eight CLDR problems suggest that the proposed methods achieve remarkable improvements in terms of Mean Reciprocal Rank compared to several baselines.
Learning Robust Search Strategies Using a Bandit-Based Approach
Effective solving of constraint problems often requires choosing good or specific search heuristics. However, choosing or designing a good search heuristic is non-trivial and is often a manual process. In this paper, rather than manually choosing/designing search heuristics, we propose the use of bandit-based learning techniques to automatically select search heuristics. Our approach is online where the solver learns and selects from a set of heuristics during search. The goal is to obtain automatic search heuristics which give robust performance. Preliminary experiments show that our adaptive technique is more robust than the original search heuristics. It can also outperform the original heuristics.
Information Retrieval System Explained Using Text Mining!
While searching for things over internet, I always wondered, what kind of algorithms might be running behind these search engines which provide us with the most relevant information? How do they decide which result to show for which set of search keywords. This might be a no brainer for a few people, but definitely an interesting problem for some of the best brains around the world. To find the answer, I read every guide, tutorial, learning material that came my way. Information retrieval system is a network of algorithms, which facilitate the search of relevant data / documents as per the user requirement.
New Google AdWords Campaigns Use Machine Learning to Maximize Conversion Value - Search Engine Journal
Google has introduced new shopping campaigns for AdWords, which utilize automation and machine learning to maximize conversion value. If an advertiser were to define their conversion value as "revenue," for example, then AdWords will automatically optimize the shopping campaign to maximize revenue based on budget constraints. Standard shopping campaigns will continue to be offered along with Google AdWords' new goal-optimized shopping campaigns. Google boasts that the new shopping campaign type "offers a fully-automated solution to drive sales and reach more customers." The new shopping campaigns will be automatically optimized to help marketers achieve their specific goal, whether it's maximizing conversion value or maximizing conversion value at a specific return on ad spend.
Private Sequential Learning
Tsitsiklis, John N., Xu, Kuang, Xu, Zhi
We formulate a private learning model to study an intrinsic tradeoff between privacy and query complexity in sequential learning. Our model involves a learner who aims to determine a scalar value, $v^*$, by sequentially querying an external database and receiving binary responses. In the meantime, an adversary observes the learner's queries, though not the responses, and tries to infer from them the value of $v^*$. The objective of the learner is to obtain an accurate estimate of $v^*$ using only a small number of queries, while simultaneously protecting her privacy by making $v^*$ provably difficult to learn for the adversary. Our main results provide tight upper and lower bounds on the learner's query complexity as a function of desired levels of privacy and estimation accuracy. We also construct explicit query strategies whose complexity is optimal up to an additive constant.
Add SEO to the List of Everything Being Transformed by Artificial Intelligence
Understanding SEO is the first rule for being able to optimize your site and its content so that would-be customers can actually find your company online. When your SEO is on point, you are more likely appear in search engine results when someone looks for the type of products or services you sell. However, to be able to truly understand SEO, you need to comprehend how modern search engines work -- and that means understanding artificial intelligence, or AI. AI is a technological advancement that enables a combination of hardware and software to function like a human brain -- minus the inherent flaws in logic and the relatively small memory capacity. It makes it possible to not only analyze large amounts of data but to draw meaningful insights about the information.
Personalizing Dialogue Agents: I have a dog, do you have pets too?
Zhang, Saizheng, Dinan, Emily, Urbanek, Jack, Szlam, Arthur, Kiela, Douwe, Weston, Jason
Chit-chat models are known to have several problems: they lack specificity, do not display a consistent personality and are often not very captivating. In this work we present the task of making chit-chat more engaging by conditioning on profile information. We collect data and train models to (i) condition on their given profile information; and (ii) information about the person they are talking to, resulting in improved dialogues, as measured by next utterance prediction. Since (ii) is initially unknown our model is trained to engage its partner with personal topics, and we show the resulting dialogue can be used to predict profile information about the interlocutors.