Goto

Collaborating Authors

 Media


Semantic Search as Extractive Paraphrase Span Detection

arXiv.org Artificial Intelligence

In this paper, we approach the problem of semantic search by framing the search task as paraphrase span detection, i.e. given a segment of text as a query phrase, the task is to identify its paraphrase in a given document, the same modelling setup as typically used in extractive question answering. On the Turku Paraphrase Corpus of 100,000 manually extracted Finnish paraphrase pairs including their original document context, we find that our paraphrase span detection model outperforms two strong retrieval baselines (lexical similarity and BERT sentence embeddings) by 31.9pp and 22.4pp respectively in terms of exact match, and by 22.3pp and 12.9pp in terms of token-level F-score. This demonstrates a strong advantage of modelling the task in terms of span retrieval, rather than sentence similarity. Additionally, we introduce a method for creating artificial paraphrase data through back-translation, suitable for languages where manually annotated paraphrase resources for training the span detection model are not available.


Josh Duggar child pornography trial: Both sides rest their case

FOX News

Fox News Flash top entertainment and celebrity headlines are here. Check out what clicked this week in entertainment. It appears Josh Duggar's child pornography trial is coming to a close. The defense rested Tuesday in the Arkansas federal trial of the former reality TV star after a prosecutor sharply questioned a computer expert during the state's cross-examination. Duggar, 33, is charged with receiving and possessing child pornography and faces up to 20 years in prison on each count if convicted.


Liquidity Group, Provides Reach Mobile with an $8 Million Funding Transaction with only 24 …

#artificialintelligence

"Our machine learning platform enabled us to deploy funding to support Reach's plans of continued growth to provide solutions to meet the needs of …


Scaling Language Models: Methods, Analysis & Insights from Training Gopher

arXiv.org Artificial Intelligence

Natural language communication is core to intelligence, as it allows ideas to be efficiently shared between humans or artificially intelligent systems. The generality of language allows us to express many intelligence tasks as taking in natural language input and producing natural language output. Autoregressive language modelling -- predicting the future of a text sequence from its past -- provides a simple yet powerful objective that admits formulation of numerous cognitive tasks. At the same time, it opens the door to plentiful training data: the internet, books, articles, code, and other writing. However this training objective is only an approximation to any specific goal or application, since we predict everything in the sequence rather than only the aspects we care about. Yet if we treat the resulting models with appropriate caution, we believe they will be a powerful tool to capture some of the richness of human intelligence. Using language models as an ingredient towards intelligence contrasts with their original application: transferring text over a limited-bandwidth communication channel. Shannon's Mathematical Theory of Communication (Shannon, 1948) linked the statistical modelling of natural language with compression, showing that measuring the cross entropy of a language model is equivalent to measuring its compression rate.


Ethical and social risks of harm from Language Models

arXiv.org Artificial Intelligence

This paper aims to help structure the risk landscape associated with large-scale Language Models (LMs). In order to foster advances in responsible innovation, an in-depth understanding of the potential risks posed by these models is needed. A wide range of established and anticipated risks are analysed in detail, drawing on multidisciplinary expertise and literature from computer science, linguistics, and social sciences. We outline six specific risk areas: I. Discrimination, Exclusion and Toxicity, II. Information Hazards, III. Misinformation Harms, V. Malicious Uses, V. Human-Computer Interaction Harms, VI. Automation, Access, and Environmental Harms. The first area concerns the perpetuation of stereotypes, unfair discrimination, exclusionary norms, toxic language, and lower performance by social group for LMs. The second focuses on risks from private data leaks or LMs correctly inferring sensitive information. The third addresses risks arising from poor, false or misleading information including in sensitive domains, and knock-on risks such as the erosion of trust in shared information. The fourth considers risks from actors who try to use LMs to cause harm. The fifth focuses on risks specific to LLMs used to underpin conversational agents that interact with human users, including unsafe use, manipulation or deception. The sixth discusses the risk of environmental harm, job automation, and other challenges that may have a disparate effect on different social groups or communities. In total, we review 21 risks in-depth. We discuss the points of origin of different risks and point to potential mitigation approaches. Lastly, we discuss organisational responsibilities in implementing mitigations, and the role of collaboration and participation. We highlight directions for further research, particularly on expanding the toolkit for assessing and evaluating the outlined risks in LMs.


Apple Music's Siri-only plan seems on track to arrive with iOS 15.2

Engadget

Apple Music's recently announced Voice Plan will launch alongside iOS 15.2, according to the patch notes the company shared for the update's release candidate. When Apple first announced the more affordable tier at its fall Mac event in October, the company said it would become available "later this fall" in 17 countries, including the US, UK and Canada. Apple also confirmed Apple Music Voice Plan will launch with iOS 15.2 pic.twitter.com/6uHeaTdr41 The plan will offer access to Apple Music's entire song catalog for $5 per month, provided you're willing to rely on Siri for control. You can play specific tracks and playlists, as well as complete albums on your Apple devices.


Nate Silver savages media study claiming harsher treatment of Biden compared to Trump: 'Complete crap'

FOX News

In media news today, CNN and Chris Cuomo issue scathing statements against each other, the former anchor announces he's leaving his SiriusXM radio show, and a New York Times op-ed gets mocked for fearing free library is contributing to gentrification. Pollster Nate Silver on Monday savaged the analytics behind a recent Washington Post column claiming President Biden was being treated just as badly, or worse, by the media than former President Trump. In the piece published last week, liberal columnist Dana Milbank complained about Biden's media coverage being overly tough and implored journalists to do "soul-searching" and "think about what it is we're delivering to people." In a series of tweets, Silver argued the piece's "sentiment analysis" measuring the positivity and negativity of particular articles written about Trump and Biden was "complete crap," and gave examples to show how the data could be skewed more positively or negatively than it should have been. "To this good thread explaining why the'sentiment analysis' cited in the [Dana Milbank] WaPo article this weekend is complete crap--the analysis was used to make the claim that the press is just negative toward Biden as Trump--I'll also add a couple of comments based on their data," Silver wrote.


Report: AI/ML impact video entertainment industry

#artificialintelligence

Often considered a'solution for everything', AI will expand its impact as the video entertainment industry realises its benefits to a variety of applications, according to a report commissioned by mobile and video technology developer InterDigital and written by market research firm Futuresource Consulting, which examines the industry influence of artificial intelligence (AI) and machine learning (ML) on applications across the video supply chain. The report, AI and Machine Learning in the Video Industry: New Opportunities for the Entertainment Sector, investigates the emerging uses of AI across segments of the media industry and highlights key examples of how AI is employed today and might develop in the future. Valued at roughly $84 billion, the global video entertainment industry is fuelled by a five-stage supply chain comprised of media creation, preparation, distribution, playout and delivery, and consumption. AI can be incorporated within several applications across the ecosystem, from encoding to transmission to decoding to post-processing. With a wide variety of applications, including auto tagging metadata, creating transcripts, conducting quality control, flagging inappropriate content, or even service personalization, AI's strength is extracting patterns from'big data' where traditional algorithms might fail.


Vewd Welcomes Intertrust ExpressPlay To Operator TV Ecosystem

#artificialintelligence

"Operator TV is an open platform to give Pay TV Operators all of the features they need to make Smart TV their domain and have an alternative to set-top boxes to control the experience on the most important screen of the home," said Marco Frattolin, Head of Operator Products at Vewd. "We're pleased that Intertrust ExpressPlay has become an Operator TV partner, enabling Pay TV Operators to select a leading broadcast and IP content security solution. Together, we can help Pay TV operators own the Smart TV." ExpressPlay XCA is a key component of the ExpressPlay Media Security Suite, which also includes a cloud-based and studio trusted multi-DRM service, comprehensive anti-piracy services, and an offline multi-DRM platform.


'He touched a nerve': how the first piece of AI music was born in 1956

The Guardian

On the evening of 9 August 1956, a couple of hundred people squeezed into a student union lounge for a concert recital at the University of Illinois Urbana-Champaign, about 130 miles outside Chicago. Student performances didn't usually attract so many people, but this was an exceptional case, the debut of the Illiac Suite: String Quartet No 4, that a member of the chemistry faculty, Lejaren Hiller Jr, had devised with the school's one and only computer, the Illiac I. Decades before today's artificial intelligence pop stars, Auto-Tune and deepfake compositions was Hiller's piece, described by the New York Times in his 1994 obituary as "the first substantial piece of music composed on a computer" – and indeed by a computer. One of the four musicians who performed the piece that night was George Andrix, a violist and composition student at the university. Now 89, Andrix remembers an auditorium packed with people "who showed up to see what this monster of a computer could do." The Illiac I, short for Illinois Automatic Computer, was the first supercomputer to be housed by an academic institution.