Goto

Collaborating Authors

 Media


How is Fatherhood Framed Online in Singapore?

arXiv.org Artificial Intelligence

The proliferation of discussion about fatherhood in Singapore attests to its significance, indicating the need for an exploration of how fatherhood is framed, aiding policy-making around fatherhood in Singapore. Sound and holistic policy around fatherhood in Singapore may reduce stigma and apprehension around being a parent, critical to improving the nation's flagging birth rate. We analyzed 15,705 articles and 56,221 posts to study how fatherhood is framed in Singapore across a range of online platforms (news outlets, parenting forums, Twitter). We used NLP techniques to understand these differences. While fatherhood was framed in a range of ways on the Singaporean online environment, it did not seem that fathers were framed as central to the Singaporean family unit. A strength of our work is how the different techniques we have applied validate each other.


Revisiting Cross-Lingual Summarization: A Corpus-based Study and A New Benchmark with Improved Annotation

arXiv.org Artificial Intelligence

Most existing cross-lingual summarization (CLS) work constructs CLS corpora by simply and directly translating pre-annotated summaries from one language to another, which can contain errors from both summarization and translation processes. To address this issue, we propose ConvSumX, a cross-lingual conversation summarization benchmark, through a new annotation schema that explicitly considers source input context. ConvSumX consists of 2 sub-tasks under different real-world scenarios, with each covering 3 language directions. We conduct thorough analysis on ConvSumX and 3 widely-used manually annotated CLS corpora and empirically find that ConvSumX is more faithful towards input text. Additionally, based on the same intuition, we propose a 2-Step method, which takes both conversation and summary as input to simulate human annotation process. Experimental results show that 2-Step method surpasses strong baselines on ConvSumX under both automatic and human evaluation. Analysis shows that both source input text and summary are crucial for modeling cross-lingual summaries.


Answering Ambiguous Questions via Iterative Prompting

arXiv.org Artificial Intelligence

In open-domain question answering, due to the ambiguity of questions, multiple plausible answers may exist. To provide feasible answers to an ambiguous question, one approach is to directly predict all valid answers, but this can struggle with balancing relevance and diversity. An alternative is to gather candidate answers and aggregate them, but this method can be computationally costly and may neglect dependencies among answers. In this paper, we present AmbigPrompt to address the imperfections of existing approaches to answering ambiguous questions. Specifically, we integrate an answering model with a prompting model in an iterative manner. The prompting model adaptively tracks the reading process and progressively triggers the answering model to compose distinct and relevant answers. Additionally, we develop a task-specific post-pretraining approach for both the answering model and the prompting model, which greatly improves the performance of our framework. Empirical studies on two commonly-used open benchmarks show that AmbigPrompt achieves state-of-the-art or competitive results while using less memory and having a lower inference latency than competing approaches. Additionally, AmbigPrompt also performs well in low-resource settings. The code are available at: https://github.com/sunnweiwei/AmbigPrompt.


Physics-Driven Diffusion Models for Impact Sound Synthesis from Videos

arXiv.org Artificial Intelligence

Modeling sounds emitted from physical object interactions is critical for immersive perceptual experiences in real and virtual worlds. Traditional methods of impact sound synthesis use physics simulation to obtain a set of physics parameters that could represent and synthesize the sound. However, they require fine details of both the object geometries and impact locations, which are rarely available in the real world and can not be applied to synthesize impact sounds from common videos. On the other hand, existing video-driven deep learning-based approaches could only capture the weak correspondence between visual content and impact sounds since they lack of physics knowledge. In this work, we propose a physics-driven diffusion model that can synthesize high-fidelity impact sound for a silent video clip. In addition to the video content, we propose to use additional physics priors to guide the impact sound synthesis procedure. The physics priors include both physics parameters that are directly estimated from noisy real-world impact sound examples without sophisticated setup and learned residual parameters that interpret the sound environment via neural networks. We further implement a novel diffusion model with specific training and inference strategies to combine physics priors and visual information for impact sound synthesis. Experimental results show that our model outperforms several existing systems in generating realistic impact sounds. More importantly, the physics-based representations are fully interpretable and transparent, thus enabling us to perform sound editing flexibly.


Multilingual Coreference Resolution in Multiparty Dialogue

arXiv.org Artificial Intelligence

Existing multiparty dialogue datasets for entity coreference resolution are nascent, and many challenges are still unaddressed. We create a large-scale dataset, Multilingual Multiparty Coref (MMC), for this task based on TV transcripts. Due to the availability of gold-quality subtitles in multiple languages, we propose reusing the annotations to create silver coreference resolution data in other languages (Chinese and Farsi) via annotation projection. On the gold (English) data, off-the-shelf models perform relatively poorly on MMC, suggesting that MMC has broader coverage of multiparty coreference than prior datasets. On the silver data, we find success both using it for data augmentation and training from scratch, which effectively simulates the zero-shot cross-lingual setting.


Meta unveils Voicebox AI: Should we all be worried?

FOX News

Meta's latest artificial intelligence model called Voicebox is a customized text to speech product that can mimic any specific voice of your choosing.


Netflix invents new green-screen filming method using magenta light

New Scientist

Netflix researchers have created a new type of AI-powered green-screen technology that can produce realistic visual effects for film and television in real time. Green-screen technology is routinely used to capture footage of actors that can then be inserted in the foreground of virtual or prerecorded scenes. To do this, actors are filmed against a bright green background, which is easily isolated and removed digitally. This process can be done automatically with reasonable accuracy, such as in television weather forecasts, but it can be thrown by items of green clothing or by transparent or fine objects, like wisps of hair. When greater accuracy is needed in films or television series, specialist operators tweak settings manually, sometimes requiring hours to perfect a shot.


New York uses drones to monitor shark activity amid rise in encounters

FOX News

George Gorman, Long Island regional director for the NY State office of Parks and Recreation, explains how drones are being used to track sharks and ensure swimmers' safety. Authorities are using drones to monitor Long Island waters following a flurry of recent incidents with sharks off New York shores. Earlier this week, five people reported being bitten by sharks at popular beaches. In response to encounters there and in other police jurisdictions, the Suffolk County Police Department said it would increase its shark patrols, using drones for an aerial view. "While residents are encouraged to enjoy the summer at the beach, swimmers should remain vigilant when in the water. If you see a shark, or a pod of bunker fish that attract the predators, calmly exit the water and alert the lifeguard on duty or a local official," the department said on Facebook.


How 'Indiana Jones and the Dial of Destiny' De-Aged Harrison Ford

WIRED

Near the end of Indiana Jones and the Dial of Destiny, Nazis attempt to pull off one of the oldest tropes in entertainment: using the movie's titular dial, the Antikythera, to travel back to 1939 and assassinate Adolf Hitler. As their Luftwaffe aircraft bears down on a time warp, the scientist Jürgen Voller (Mads Mikkelsen), who hopes to install himself as the führer and win the war, turns to Indiana Jones and demands he witness "history's greatest moment--its end." To enter the past, then, is to end history. It's Voller's motto, but also the movie's--a nod to the de-aging technology that has made it possible. Thanks to several tools--AI, CGI, other acronyms--80-year-old Harrison Ford spends roughly 25 minutes of the film looking like the Indiana Jones of the early 1980s.


It's been 100 days since American journalist detained in Russia, model returns to court and more top headlines

FOX News

Subscribe now to get Fox News First in your email. And here's what you need to know to start your day ... 100 DAYS - Today marks 100 days since WSJ reporter Evan Gershkovich was detained by Russia. Journalism is not a crime, and we will not rest until Evan is released. BACK IN COURT – OnlyFans model Courtney Clenney, held in prison without bail for the fatal stabbing of her live-in boyfriend Christian Obumseli in April 2022, returns to a Florida courtroom Friday. DESPERATE DODGING - Experts astonished White House invokes Hatch Act to avoid Hunter Biden cocaine question.