SPE
The Stanford Natural Language Processing Group
Tokenization of raw text is a standard pre-processing step for many NLP tasks. For English, tokenization usually involves punctuation splitting and separation of some affixes like possessives. Other languages require more extensive token pre-processing, which is usually called segmentation. The Stanford Word Segmenter currently supports Arabic and Chinese. The provided segmentation schemes have been found to work well for a variety of applications.
Google sets the bar high for its Oct. phone reveal
Google has helped build intense speculation for its October 4 event in San Francisco, where it's expected to reveal new phones aimed at consumers that will power a new virtual reality platform, and possibly other smart home devices. Now that the buzz has reached a football-stadium roar, here comes the hard part: living up to the hype. Google has been teasing the event as one for the history books. A tweet Monday from Hiroshi Lockheimer, the company's senior vice president of Android, Chrome OS and Google Play, turned up the volume on the buzz. We announced the 1st version of Android 8 years ago today.
Google Translate Just Got 60% Better By Working on Whole Sentences
If you are translating text or speech, it seems obvious that you should read a whole sentence before figuring out what it means. But this hasn't been so easy for computers--in part because the work sucks up so many resources. So Google Translate has had to get by with looking at pieces of sentences, words, and phrases, and translating them individually. On Tuesday, Google announced a new system called Google Neural Machine Translation (GNMT) that works on whole sentences and improves accuracy about 60% on average over the old phrase-based machine translation (PBMT), including on notoriously difficult Chinese-to-English translations. This is the first translation to roll out to the Google Translate mobile and web apps, available now.
Google Translate Receives Huge Accuracy Boost Through The Power Of Neural Networks
When Google launched Google Translate 10 years ago, the key algorithm behind the service was Phrase-Based Machine Translation. The translations provided by the service have since vastly improved due to the developments in machine intelligence, but the recent addition of neural networks has provided Google Translate with the biggest boost that it has ever received. Language is naturally phrase-based, which is why translating between languages is not as simple as plugging in the translation of words in sentences. While computers have been developed to handle phrase-based translation, there are still nuances in languages that the machines are not able to understand. Google has now deployed the Google Neural Machine Translation system, which utilizes machine learning and neural networks to provide a massive boost in translation accuracy.
Dimension Reduction and Intuitive Feature Engineering for Machine Learning
In the previous parts of this series, we looked at an overview of some popular tricks for feature engineering, and examined those tricks in greater detail. In this part, we continue our closer examination of these approaches with a deeper dive into the final techniques described in Part 1. The examples discussed in this article can be reproduced with the source code and datasets available here. As an analyst, you savor the scenario in which you have a lot of data. But, with a lot of data comes the added complexity of analyzing and making better sense of that data.
Machine Learning and CDS Transparency
One of the many questions in the design and use of Clinical Decision Support software is whether or not the user can recreate the logic used by the system in reaching its conclusions and recommendationsโor alerts, or suggestions. If the CDS is based on sound medical logic, perhaps supported by specific reference material, then the user could in principle reach the same conclusions by reading the same literature, or perhaps reach a different conclusion. This transparency was part of the proposed criteria for some CDS systems not falling under FDA regulation in 2015 federal draft legislation--which didn't pass. The FDA has otherwise not been forthcoming on the general subject of CDS despite many pleas for guidance, and a draft guidance in this domain is an as yet unfulfilled part of the 2015 strategic plan. However underlying logic and science is not the only way to build "artificial intelligence" (AI), which might in some instances turn out to be artificial mediocrity if not artificial stupidity.
Why data is the new coal
"Is data the new oil?" asked proponents of big data back in 2012 in Forbes magazine. By 2016, and the rise of big data's turbo-powered cousin deep learning, we had become more certain: "Data is the new oil," stated Fortune. Amazon's Neil Lawrence has a slightly different analogy: Data, he says, is coal. Not coal today, though, but coal in the early days of the 18th century, when Thomas Newcomen invented the steam engine. A Devonian ironmonger, Newcomen built his device to pump water out of the south west's prolific tin mines. The problem, as Lawrence told the Re-Work conference on Deep Learning in London, was that the pump was rather more useful to those who had a lot of coal than those who didn't: it was good, but not good enough to buy coal in to run it.
The Good and Bad of Microsoft's Cloud Strategy
Microsoft is well positioned to give Amazon a run for its money in the cloud market, but it needs to break away from its Microsoft-centric approach. Seeing as how the cloud has been tied to digital transformation, and seeing as how more businesses are embarking on digital transformation projects, it makes perfect sense to me that cloud has been one of the hot topics at Microsoft's Ignite conference for enterprise IT, taking place this week in Atlanta. Microsoft has an interesting position in cloud, in that it was simultaneously early and late to the market. Almost 20 years ago, Microsoft launched Bing, and to support it, the company had to build out a massively scalable, global cloud network. Google had done this with its search platform, and Amazon had done similar to support its e-commerce business.
Keeping AI Well Behaved: How Do We Engineer An Artificial System That Has Values?
Imagine you're sitting in a self-driving car that's about to make a left turn into on-coming traffic. One small AI system in the car will be responsible for making the vehicle turn, one system might speed it up or hit the brakes, other systems will have sensors that detect obstacles, and yet another system may be in communication with other vehicles on the road. Each system has its own goals -- starting or stopping, turning or traveling straight, recognizing potential problems, etc. -- but they also have to all work together toward one common goal: turning into traffic without causing an accident. Harvard professor and Future of Life researcher, David Parkes, is trying to solve just this type of problem. Parkes told FLI, "The particular question I'm asking is: If we have a system of AIs, how can we construct rewards for individual AIs, such that the combined system is well behaved?"
Artificial Intelligence - Applications in Insurance Industry
Over the past few years, Artificial Intelligence (AI) as a technology has matured and come into its own. With each passing day, experts across industries identify yet another AI application that has the potential to change millions of human lives. California-based University of Southern California's Viterbi School of Engineering and its School of Social Work recently announced that they had joined forces to launch the Center on Artificial Intelligence for Social Solutions. Also in California, University of California Berkeley unveiled its Center for Human-Compatible Artificial Intelligence. Google's DeepMind is learning how to better apply radiotherapy to cancer patients to reduce the impact of dangerous doses of radiation on areas surrounding a tumor.