Media
ProsAudit, a prosodic benchmark for self-supervised speech models
de Seyssel, Maureen, Lavechin, Marvin, Titeux, Hadrien, Thomas, Arthur, Virlet, Gwendal, Revilla, Andrea Santos, Wisniewski, Guillaume, Ludusan, Bogdan, Dupoux, Emmanuel
We present ProsAudit, a benchmark in English to assess structural prosodic knowledge in self-supervised learning (SSL) speech models. It consists of two subtasks, their corresponding metrics, and an evaluation dataset. In the protosyntax task, the model must correctly identify strong versus weak prosodic boundaries. In the lexical task, the model needs to correctly distinguish between pauses inserted between words and within words. We also provide human evaluation scores on this benchmark. We evaluated a series of SSL models and found that they were all able to perform above chance on both tasks, even when evaluated on an unseen language. However, non-native models performed significantly worse than native ones on the lexical task, highlighting the importance of lexical knowledge in this task. We also found a clear effect of size with models trained on more data performing better in the two subtasks.
Automatic Creation of Named Entity Recognition Datasets by Querying Phrase Representations
Kim, Hyunjae, Yoo, Jaehyo, Yoon, Seunghyun, Kang, Jaewoo
Most weakly supervised named entity recognition (NER) models rely on domain-specific dictionaries provided by experts. This approach is infeasible in many domains where dictionaries do not exist. While a phrase retrieval model was used to construct pseudo-dictionaries with entities retrieved from Wikipedia automatically in a recent study, these dictionaries often have limited coverage because the retriever is likely to retrieve popular entities rather than rare ones. In this study, we present a novel framework, HighGEN, that generates NER datasets with high-coverage pseudo-dictionaries. Specifically, we create entity-rich dictionaries with a novel search method, called phrase embedding search, which encourages the retriever to search a space densely populated with various entities. In addition, we use a new verification process based on the embedding distance between candidate entity mentions and entity types to reduce the false-positive noise in weak labels generated by high-coverage dictionaries. We demonstrate that HighGEN outperforms the previous best model by an average F1 score of 4.7 across five NER benchmark datasets.
Drone footage shows shark circling man and small child at Alabama beach
The Gulf of Mexico has around 50 species of sharks, with around 20 to 30 species that beachgoers and fishermen can encounter. A shark was captured on drone footage Monday circling a man and a child swimming at a popular beach in Alabama. The footage, taken by 15-year-old Jackson Silvio and obtained by Fox News Digital, shows the man and child wading further out into the water at Orange Beach. At one point the shark appeared to swim just within a few feet of the man. The shark can be seen following them, swimming in a circle as it gets closer.
The 20 most puzzling questions in modern life revealed - so do YOU know the answers?
What is an NFT? (34%) Non-fungible tokens (NFTs) are generally digital art pieces or music that can be bought or traded online. These are unique computer files encrypted with an artist's signature. As a result, they cannot be replicated, acting as a digital certificate of ownership and authenticity. In other words, buying an NFT is almost like the more traditional purchasing of fine art - except in a digital form. Artists can sell pieces that may be tricky to advertise otherwise, such as digital stickers.
Amazon Echo Pop Review (2023): Fun To Look At
Its newest player has arrived: the Echo Pop, a $40 smart speaker with a half-circle body instead of the fully rounded forms of Amazon's other Echo speakers. I've been enjoying the Echo Pop as a small desk companion. Its sound is fine enough for its price--though other, similarly priced Amazon speakers will be a better music experience--but the biggest appeal of the Echo Pop is easily the fun colors and interesting form factor it has. When it's next to the Echo Dot (5th Gen), the Echo Pop looks to be the same size--that is, if you're looking head-on at the Pop's flat, circular face. But when you take a look at the Echo Pop from the side, you can see it's only about two-thirds as deep. Where you really can see the size difference is on the bottom.
Amazon is making a HUGE change to Alexa's voice - here's what it means for your smart assistant
Amazon has revealed a huge change that will make interacting with its smart speakers a lot less fun. The tech giant is retiring all three celebrity voices for its smart speakers โ Samuel L. Jackson, Shaquille O'Neal and Melissa McCarthy. Amazon offered the superstar voices for $4.99 each as an alternative to Alexa, but these are no longer available for purchase on its website. Amazon, which released its fifth generation Echo Dot smart speaker last year, said customers can contact them for a refund. The feature was for US users only, although the tech giant does offer alternative voices for its smart assistant in the UK, such as Santa Claus.
Next generation arms race could cause 'extinction' event akin to nuclear war, pandemic: tech chief
Artificial intelligence could lead to extinction and should be a global priority on the scale of nuclear war and pandemics, Center for AI Safety chief Dan Hendrycks said. An artificial intelligence arms race between countries and corporations to see who can develop the most powerful AI machines could create an existential threat to humanity, the co-founder of an AI safety nonprofit told Fox News. "AI could pose the risk of extinction, and part of the reason for this is because we're currently locked in an AI arms race," Center for AI Safety Executive Director Dan Hendrycks said. "We're building increasingly powerful technologies, and we don't know how to completely control them or understand them." Sam Altman, CEO of OpenAI, signed the Center for AI Safety's statement saying that AI poses an existential threat to humanity.
Graph Neural Tangent Kernel: Convergence on Large Graphs
Krishnagopal, Sanjukta, Ruiz, Luana
Graph neural networks (GNNs) achieve remarkable performance in graph machine learning tasks but can be hard to train on large-graph data, where their learning dynamics are not well understood. We investigate the training dynamics of large-graph GNNs using graph neural tangent kernels (GNTKs) and graphons. In the limit of large width, optimization of an overparametrized NN is equivalent to kernel regression on the NTK. Here, we investigate how the GNTK evolves as another independent dimension is varied: the graph size. We use graphons to define limit objects -- graphon NNs for GNNs, and graphon NTKs for GNTKs -- , and prove that, on a sequence of graphs, the GNTKs converge to the graphon NTK. We further prove that the spectrum of the GNTK, which is related to the directions of fastest learning which becomes relevant during early stopping, converges to the spectrum of the graphon NTK. This implies that in the large-graph limit, the GNTK fitted on a graph of moderate size can be used to solve the same task on the large graph, and to infer the learning dynamics of the large-graph GNN. These results are verified empirically on node regression and classification tasks.
The Stable Artist: Steering Semantics in Diffusion Latent Space
Brack, Manuel, Schramowski, Patrick, Friedrich, Felix, Hintersdorf, Dominik, Kersting, Kristian
Large, text-conditioned generative diffusion models have recently gained a lot of attention for their impressive performance in generating high-fidelity images from text alone. However, achieving high-quality results is almost unfeasible in a one-shot fashion. On the contrary, text-guided image generation involves the user making many slight changes to inputs in order to iteratively carve out the envisioned image. However, slight changes to the input prompt often lead to entirely different images being generated, and thus the control of the artist is limited in its granularity. To provide flexibility, we present the Stable Artist, an image editing approach enabling fine-grained control of the image generation process. The main component is semantic guidance (SEGA) which steers the diffusion process along variable numbers of semantic directions. This allows for subtle edits to images, changes in composition and style, as well as optimization of the overall artistic conception. Furthermore, SEGA enables probing of latent spaces to gain insights into the representation of concepts learned by the model, even complex ones such as 'carbon emission'. We demonstrate the Stable Artist on several tasks, showcasing high-quality image editing and composition.
Detection of Late Blight Disease in Tomato Leaf Using Image Processing Techniques
Farooq, Muhammad Shoaib, Arif, Tabir, Riaz, Shamyla
=One of the most frequently farmed crops is the tomato crop. Late blight is the most prevalent tomato disease in the world, and often causes a significant reduction in the production of tomato crops. The importance of tomatoes as an agricultural product necessitates early detection of late blight. It is produced by the fungus Phytophthora. The earliest signs of late blight on tomatoes are unevenly formed, water-soaked lesions on the leaves located on the plant canopy's younger leave White cottony growth may appear in humid environments evident on the undersides of the leaves that have been impacted. Lesions increase as the disease proceeds, turning the leaves brown to shrivel up and die. Using picture segmentation and the Multi-class SVM technique, late blight disorder is discovered in this work. Image segmentation is employed for separating damaged areas on leaves, and the Multi-class SVM method is used for reliable disease categorization. 30 reputable studies were chosen from a total of 2770 recognized papers. The primary goal of this study is to compile cutting-edge research that identifies current research trends, problems, and prospects for late blight detection. It also looks at current approaches for applying image processing to diagnose and detect late blight. A suggested taxonomy for late blight detection has also been provided. In the same way, a model for the development of the solutions to problems is also presented. Finally, the research gaps have been presented in terms of open issues for the provision of future directions in image processing for the researchers.