Goto

Collaborating Authors

 Information Retrieval


Google employees plan walkout over censored Chinese search engine

Engadget

Just weeks after Google employees walked out of offices to protest the way the company dealt with claims of sexual misconduct, Google is bracing itself for another worldwide protest. This time, it's over Google's ominous Project Dragonfly, and human rights organization Amnesty International is throwing its whole weight behind it. Project Dragonfly has already received global backlash, with Google employees themselves calling publicly for an ethics review into the proposed censored Chinese search engine. According to leaked documents, the search app will automatically identify websites blocked by China's so-called Great Firewall. This includes information on free speech, current affairs and political opposition, plus historical references to specific events (such as the 1989 Tiananmen Square massacre) and books that negatively feature authoritarian governments.


Google employees sign letter against censored search engine for China

The Guardian

A group of Google employees published an open letter on Tuesday calling on their employer to cancel its plans to build a censored search engine for China, the latest expression of worker unrest at a company that earlier this month saw thousands stage walkouts over its handling of sexual misconduct cases. Google's plan for returning to China, which is known as Project Dragonfly and would reportedly allow the Chinese government to blacklist certain search terms and control air quality data, has garnered significant backlash internally since it was first reported on in August. More than 1,400 Google employees signed an internal petition criticizing the lack of transparency around the project, and at least one employee resigned in protest. But Tuesday's letter, which was initially signed by nine current Google employees, is a bold step for employees of a company that prizes internal transparency but considers leaking information to be not "Googley". Organizers of the letter said they would continuously update the letter as more employees signed on; by midday there were more than 50 signers.


$HS^2$: Active Learning over Hypergraphs

arXiv.org Machine Learning

We propose a hypergraph-based active learning scheme which we term $HS^2$, $HS^2$ generalizes the previously reported algorithm $S^2$ originally proposed for graph-based active learning with pointwise queries [Dasarathy et al., COLT 2015]. Our $HS^2$ method can accommodate hypergraph structures and allows one to ask both pointwise queries and pairwise queries. Based on a novel parametric system particularly designed for hypergraphs, we derive theoretical results on the query complexity of $HS^2$ for the above described generalized settings. Both the theoretical and empirical results show that $HS^2$ requires a significantly fewer number of queries than $S^2$ when one uses $S^2$ over a graph obtained from the corresponding hypergraph via clique expansion.


The Use of NLP to Extract Unstructured Medical Data From Text - insideBIGDATA

#artificialintelligence

When working in healthcare, a lot of the relevant information for making accurate predictions and recommendations is only available in free-text clinical notes. Much of this data is trapped in free-text documents in unstructured form. This data is needed in order to make healthcare decisions. Hence, it is important to be able to extract data in the best possible way such that the information obtained can be analyzed and used. State-of-the-art NLP algorithms can extract clinical data from text using deep learning techniques such as healthcare-specific word embeddings, named entity recognition models, and entity resolution models.


Can China's new AI news anchors give Anderson Cooper a run for his money?

#artificialintelligence

China's state-owned Xinhua News Agency introduced so-called "composite anchors" on Wednesday, combining the images and voices of human anchors with artificial intelligence (AI) technology. The new AI anchors, launched by Xinhua and Beijing-based search engine operator Sogou during the World Internet Conference in Wuzhen, can deliver the news with "the same effect" as human anchors because the machine learning programme is able to synthesise realistic-looking speech, lip movements and facial expressions, according to a Xinhua news report on Wednesday. "AI anchors have officially become members of the Xinhua News Agency reporting team. They will work with other anchors to bring you authoritative, timely and accurate news information in both Chinese and English," Xinhua said. The AI anchors are now available throughout Xinhua's internet and mobile platforms such as its official Chinese and English apps, WeChat public account, and online TV webpage.


Analysis of Google's New Schema Speakable Markup - Search Engine Journal

#artificialintelligence

Google announced official support for the Schema.org The speakable specification will help Google Assistant and Google Home choose which content to read aloud. This new structured data markup is important because it may point to what you'll need to know to get more traffic should/when Google expands this structured data to all websites. The support for this new markup is currently limited to News content. However, it is likely that support for the speakable attribute will inevitably expand as Google gains experience with this new structured data markup.


Satyam: Democratizing Groundtruth for Machine Vision

arXiv.org Machine Learning

The democratization of machine learning (ML) has led to ML-based machine vision systems for autonomous driving, traffic monitoring, and video surveillance. However, true democratization cannot be achieved without greatly simplifying the process of collecting groundtruth for training and testing these systems. This groundtruth collection is necessary to ensure good performance under varying conditions. In this paper, we present the design and evaluation of Satyam, a first-of-its-kind system that enables a layperson to launch groundtruth collection tasks for machine vision with minimal effort. Satyam leverages a crowdtasking platform, Amazon Mechanical Turk, and automates several challenging aspects of groundtruth collection: creating and launching of custom web-UI tasks for obtaining the desired groundtruth, controlling result quality in the face of spammers and untrained workers, adapting prices to match task complexity, filtering spammers and workers with poor performance, and processing worker payments. We validate Satyam using several popular benchmark vision datasets, and demonstrate that groundtruth obtained by Satyam is comparable to that obtained from trained experts and provides matching ML performance when used for training.


Experimentation & Measurement for Search Engine Optimization

#artificialintelligence

For many of our potential guests, planning a trip starts at the search engine. At Airbnb, we want our product to be painless to find for past guests, and easy to discover for new ones. Search engine optimization (SEO) is the process of improving our site -- and more specifically our landing pages--to ensure that when a traveller looks for accommodations for their next trip, Airbnb is one of the top results on their favorite search engine. Search engines such as Google, Yahoo, Naver, and Baidu deploy their own fleet of "bots" across the internet to build map of the web and scrape information, or "index", from the pages that they hit. When indexing pages and ranking them for specific search queries, search engines will take into account a variety of factors, including relevance, site performance, and authority.


SimplerVoice: A Key Message & Visual Description Generator System for Illiteracy

arXiv.org Artificial Intelligence

We introduce SimplerVoice: a key message and visual description generator system to help low-literate adults navigate the information-dense world with confidence, on their own. SimplerVoice can automatically generate sensible sentences describing an unknown object, extract semantic meanings of the object usage in the form of a query string, then, represent the string as multiple types of visual guidance (pictures, pictographs, etc.). We demonstrate SimplerVoice system in a case study of generating grocery products' manuals through a mobile application. To evaluate, we conducted a user study on SimplerVoice's generated description in comparison to the information interpreted by users from other methods: the original product package and search engines' top result, in which SimplerVoice achieved the highest performance score: 4.82 on 5-point mean opinion score scale. Our result shows that SimplerVoice is able to provide low-literate end-users with simple yet informative components to help them understand how to use the grocery products, and that the system may potentially provide benefits in other real-world use cases.


Learning to Rank Query Graphs for Complex Question Answering over Knowledge Graphs

arXiv.org Artificial Intelligence

In this paper, we conduct an empirical investigation of neural query graph ranking approaches for the task of complex question answering over knowledge graphs. We experiment with six different ranking models and propose a novel self-attention based slot matching model which exploits the inherent structure of query graphs, our logical form of choice. Our proposed model generally outperforms the other models on two QA datasets over the DBpedia knowledge graph, evaluated in different settings. In addition, we show that transfer learning from the larger of those QA datasets to the smaller dataset yields substantial improvements, effectively offsetting the general lack of training data.