Information Retrieval
Google to 'shut down plans' for censored Chinese search engine
Google has been forced to abandon its specialist Chinese search engine that censors results in line with the strict government, reports have claimed. The firm is believed to have shut down an internal data analysis system which was being used to develop the search engine, known as Dragonfly. According to a report from The Intercept, this has'effectively ended' the entire project. Members of Google's privacy team raised concerns about the project back in August and it is now extremely unlikely the search engine can be built without the system, according to sources close to the project. Google has been forced to abandon its plan to launch a specialist Chinese search engine that censors results in line with the strict government.
Efficient Autotuning of Hyperparameters in Approximate Nearest Neighbor Search
Jรครคsaari, Elias, Hyvรถnen, Ville, Roos, Teemu
Approximate nearest neighbor algorithms are used to speed up nearest neighbor search in a wide array of applications. However, current indexing methods feature several hyperparameters that need to be tuned to reach an acceptable accuracy--speed trade-off. A grid search in the parameter space is often impractically slow due to a time-consuming index-building procedure. Therefore, we propose an algorithm for automatically tuning the hyperparameters of indexing methods based on randomized space-partitioning trees. In particular, we present results using randomized k-d trees, random projection trees and randomized PCA trees. The tuning algorithm adds minimal overhead to the index-building process but is able to find the optimal hyperparameters accurately. We demonstrate that the algorithm is significantly faster than existing approaches, and that the indexing methods used are competitive with the state-of-the-art methods in query time while being faster to build.
Google's China search engine project 'effectively ended': report
Members of the House Judiciary Committee peppered the head of Google about potential bias against conservatives and Russian influence and misinformation; Gillian Turner reports. Google has been forced to shut down and "effectively end" its controversial China search engine project, code-named Project Dragonfly, after members of the company's privacy team raised complaints, according to a new report. The tech giant led by CEO Sundar Pichai was forced to close a data analysis system it was using for the controversial project, according to The Intercept, citing two sources familiar with the matter. The news outlet originally broke the news that Google had been considering launching the app-based search engine. Google has not yet responded to a request for comment from Fox News.
Google CEO Sundar Pichai refuses to rule out censored Chinese search engine
Google's chief executive, Sundar Pichai, testified before the House judiciary committee on Tuesday morning, three months after his company thumbed its nose at Congress by failing to appear alongside Facebook and Twitter at a Senate hearing on election interference. In a hearing heavy on partisan theatrics, Pichai notably refused to rule out launching a censored search engine in China, a controversial plan that has garnered significant criticism from human rights organizations as well as rank-and-file Google employees. "Right now there are no plans to launch search in China," Pichai said numerous times, repeating a talking point that the company has relied on since news of the project leaked in August. Pichai characterized the Chinese search product as an "internal effort" and said the company would be "transparent" and consult with policy makers before launching in China. Pressed to rule out launching a tool that would enable censorship and surveillance in China, however, Pichai appeared to offer the company's probable justification for reentering a market that it left in 2010: "We think it's in our duty to explore possibilities to give users access to information."
Google has 'no plans' to launch Chinese search engine -CEO
Google has'no plans' to relaunch a search engine in China though it is continuing to study the idea, Chief Executive Sundar Pichai told a U.S. congressional panel on Tuesday amid increased scrutiny of big tech firms. Lawmakers and Google employees have raised concerns the company would comply with China's internet censorship and surveillance policies if it re-enters the Asian nation's search engine market. Google's main search platform has been blocked in China since 2010, but the Alphabet Inc unit has been attempting to make new inroads into the country, which has the world's largest number of smartphone users. Chief Executive Sundar Pichai told a U.S. congressional panel Google had over 100 people working on the project at one point. 'Right now, there are no plans to launch search in China,' Pichai told the U.S. House of Representatives Judiciary Committee.
Detecting weak and strong Islamophobic hate speech on social media
Islamophobic hate speech on social media inflicts considerable harm on both targeted individuals and wider society, and also risks reputational damage for the host platforms. Accordingly, there is a pressing need for robust tools to detect and classify Islamophobic hate speech at scale. Previous research has largely approached the detection of Islamophobic hate speech on social media as a binary task. However, the varied nature of Islamophobia means that this is often inappropriate for both theoretically-informed social science and effectively monitoring social media. Drawing on in-depth conceptual work we build a multi-class classifier which distinguishes between non-Islamophobic, weak Islamophobic and strong Islamophobic content. Accuracy is 77.6% and balanced accuracy is 83%. We apply the classifier to a dataset of 109,488 tweets produced by far right Twitter accounts during 2017. Whilst most tweets are not Islamophobic, weak Islamophobia is considerably more prevalent (36,963 tweets) than strong (14,895 tweets). Our main input feature is a gloVe word embeddings model trained on a newly collected corpus of 140 million tweets. It outperforms a generic word embeddings model by 5.9 percentage points, demonstrating the importan4ce of context. Unexpectedly, we also find that a one-against-one multi class SVM outperforms a deep learning algorithm.
Congress grills Google CEO over Chinese search engine plans
If you were hoping that Google chief Sundar Pichai would shed more light on his company's potential censored search engine for China... well, you'll mostly be disappointed. Rhode Island Representative David Cicilline grilled Pichai on the recently acknowledged Dragonfly project and mostly encountered attempts to downplay the significance of the engine. The Google exec stressed there were "no plans" to launch a search engine for China, and that Dragonfly was an "internal effort" and "limited" in scope. Pichai added that Google was "currently not in discussions" with Chinese officials. He also provided a non-committal answer when asked if Google would promise not to create a tool enabling Chinese surveillance.
Interval type-2 Beta Fuzzy Near set based approach to content based image retrieval
Ghozzi, Yosr, Baklouti, Nesrine, Hagras, Hani, Ayed, Mounir Ben, Alimi, Adel M.
Abstract-- In an automated search system, similarity is a key concept in solving a human task. Indeed, human process is usually a natural categorization that underlies many natural abilities such as image recovery, language comprehension, decision making, or pattern recognition. In the image search axis, there are several ways to measure the similarity between images in an image database, to a query image. Image search by content is based on the similarity of the visual characteristics of the images. The distance function used to evaluate the similarity between images depends on the criteria of the search but also on the representation of the characteristics of the image; this is the main idea of the near and fuzzy sets approaches. In this article, we introduce a new category of beta type-2 fuzzy sets for the description of image characteristics as well as the near sets approach for image recovery. Finally, we illustrate our work with examples of image recovery problems used in the real world. I. INTRODUCTION He number of daily-generated images by websites and personal archives are constantly growing. Indeed, the effective management of the rapid expansion of visual information has become a major problem and a necessity for strengthening visual search technique based on visual content [3]. This necessity is behind the emergence of new visual search techniques based on visual content. It has been widely identified that the most efficient and intuitive way to research visual information is based on the properties that are extracted from the images themselves. Researchers from different communities ("Computer Vision" [4], "Database Management", "Man-machine Interface", "Information Retrieval") were attracted by this field. Since then, the search for images by content has developed quite rapidly. The intuitive idea of "any system that analyzes or automatically organizes a set of data or knowledge must use, in one form or another, a similarity operator whose purpose is to establish similarities or the relationships that exist between the manipulated information".
Taking the Scenic Route: Automatic Exploration for Videogames
Zhan, Zeping, Aytemiz, Batu, Smith, Adam M.
Machine playtesting tools and game moment search engines require exposure to the diversity of a game's state space if they are to report on or index the most interesting moments of possible play. Meanwhile, mobile app distribution services would like to quickly determine if a freshly-uploaded game is fit to be published. Having access to a semantic map of reachable states in the game would enable efficient inference in these applications. However, human gameplay data is expensive to acquire relative to the coverage of a game that it provides. We show that off-the-shelf automatic exploration strategies can explore with an effectiveness comparable to human gameplay on the same timescale. We contribute generic methods for quantifying exploration quality as a function of time and demonstrate our metric on several elementary techniques and human players on a collection of commercial games sampled from multiple game platforms (from Atari 2600 to Nintendo 64). Emphasizing the diversity of states reached and the semantic map extracted, this work makes productive contrast with the focus on finding a behavior policy or optimizing game score used in most automatic game playing research.
Improving Similarity Search with High-dimensional Locality-sensitive Hashing
Sharma, Jaiyam, Navlakha, Saket
We propose a new class of data-independent locality-sensitive hashing (LSH) algorithms based on the fruit fly olfactory circuit. The fundamental difference of this approach is that, instead of assigning hashes as dense points in a low dimensional space, hashes are assigned in a high dimensional space, which enhances their separability. We show theoretically and empirically that this new family of hash functions is locality-sensitive and preserves rank similarity for inputs in any `p space. We then analyze different variations on this strategy and show empirically that they outperform existing LSH methods for nearest-neighbors search on six benchmark datasets. Finally, we propose a multi-probe version of our algorithm that achieves higher performance for the same query time, or conversely, that maintains performance of prior approaches while taking significantly less indexing time and memory. Overall, our approach leverages the advantages of separability provided by high-dimensional spaces, while still remaining computationally efficient