Goto

Collaborating Authors

 Media


Relational Extraction on Wikipedia Tables using Convolutional and Memory Networks

arXiv.org Artificial Intelligence

Relation extraction (RE) is the task of extracting relations between entities in text. Most RE methods extract relations from free-form running text and leave out other rich data sources, such as tables. We explore RE from the perspective of applying neural methods on tabularly organized data. We introduce a new model consisting of Convolutional Neural Network (CNN) and Bidirectional-Long Short Term Memory (BiLSTM) network to encode entities and learn dependencies among them, respectively. We evaluate our model on a large and recent dataset and compare results with previous neural methods. Experimental results show that our model consistently outperforms the previous model for the task of relation extraction on tabular data. We perform comprehensive error analyses and ablation study to show the contribution of various components of our model. Finally, we discuss the usefulness and trade-offs of our approach, and provide suggestions for fostering further research.


Speech Diarization and ASR with GMM

arXiv.org Artificial Intelligence

In this research paper, we delve into the topics of Speech Diarization and Automatic Speech Recognition (ASR). Speech diarization involves the separation of individual speakers within an audio stream. By employing the ASR transcript, the diarization process aims to segregate each speaker's utterances, grouping them based on their unique audio characteristics. On the other hand, Automatic Speech Recognition refers to the capability of a machine or program to identify and convert spoken words and phrases into a machine-readable format. In our speech diarization approach, we utilize the Gaussian Mixer Model (GMM) to represent speech segments. The inter-cluster distance is computed based on the GMM parameters, and the distance threshold serves as the stopping criterion. ASR entails the conversion of an unknown speech waveform into a corresponding written transcription. The speech signal is analyzed using synchronized algorithms, taking into account the pitch frequency. Our primary objective typically revolves around developing a model that minimizes the Word Error Rate (WER) metric during speech transcription.


A data science and machine learning approach to continuous analysis of Shakespeare's plays

arXiv.org Artificial Intelligence

The availability of quantitative text analysis methods has provided new ways of analyzing literature in a manner that was not available in the pre-information era. Here we apply comprehensive machine learning analysis to the work of William Shakespeare. The analysis shows clear changes in the style of writing over time, with the most significant changes in the sentence length, frequency of adjectives and adverbs, and the sentiments expressed in the text. Applying machine learning to make a stylometric prediction of the year of the play shows a Pearson correlation of 0.71 between the actual and predicted year, indicating that Shakespeare's writing style as reflected by the quantitative measurements changed over time. Additionally, it shows that the stylometrics of some of the plays is more similar to plays written either before or after the year they were written. For instance, Romeo and Juliet is dated 1596, but is more similar in stylometrics to plays written by Shakespeare after 1600. The source code for the analysis is available for free download. INTRODUCTION Being one of the most in influential authors in history, the analysis of the stylometrics of William Shakespeare has been a topic of substantial interest.


3D detection of roof sections from a single satellite image and application to LOD2-building reconstruction

arXiv.org Artificial Intelligence

Reconstructing urban areas in 3D out of satellite raster images has been a long-standing and challenging goal of both academical and industrial research. The rare methods today achieving this objective at a Level Of Details $2$ rely on procedural approaches based on geometry, and need stereo images and/or LIDAR data as input. We here propose a method for urban 3D reconstruction named KIBS(\textit{Keypoints Inference By Segmentation}), which comprises two novel features: i) a full deep learning approach for the 3D detection of the roof sections, and ii) only one single (non-orthogonal) satellite raster image as model input. This is achieved in two steps: i) by a Mask R-CNN model performing a 2D segmentation of the buildings' roof sections, and after blending these latter segmented pixels within the RGB satellite raster image, ii) by another identical Mask R-CNN model inferring the heights-to-ground of the roof sections' corners via panoptic segmentation, unto full 3D reconstruction of the buildings and city. We demonstrate the potential of the KIBS method by reconstructing different urban areas in a few minutes, with a Jaccard index for the 2D segmentation of individual roof sections of $88.55\%$ and $75.21\%$ on our two data sets resp., and a height's mean error of such correctly segmented pixels for the 3D reconstruction of $1.60$ m and $2.06$ m on our two data sets resp., hence within the LOD2 precision range.


ProgGP: From GuitarPro Tablature Neural Generation To Progressive Metal Production

arXiv.org Artificial Intelligence

Recent work in the field of symbolic music generation has shown value in using a tokenization based on the GuitarPro format, a symbolic representation supporting guitar expressive attributes, as an input and output representation. We extend this work by fine-tuning a pre-trained Transformer model on ProgGP, a custom dataset of 173 progressive metal songs, for the purposes of creating compositions from that genre through a human-AI partnership. Our model is able to generate multiple guitar, bass guitar, drums, piano and orchestral parts. We examine the validity of the generated music using a mixed methods approach by combining quantitative analyses following a computational musicology paradigm and qualitative analyses following a practice-based research paradigm. Finally, we demonstrate the value of the model by using it as a tool to create a progressive metal song, fully produced and mixed by a human metal producer based on AI-generated music.


On the Effectiveness of Speech Self-supervised Learning for Music

arXiv.org Artificial Intelligence

Self-supervised learning (SSL) has shown promising results in various speech and natural language processing applications. However, its efficacy in music information retrieval (MIR) still remains largely unexplored. While previous SSL models pre-trained on music recordings may have been mostly closed-sourced, recent speech models such as wav2vec2.0 have shown promise in music modelling. Nevertheless, research exploring the effectiveness of applying speech SSL models to music recordings has been limited. We explore the music adaption of SSL with two distinctive speech-related models, data2vec1.0 and Hubert, and refer to them as music2vec and musicHuBERT, respectively. We train $12$ SSL models with 95M parameters under various pre-training configurations and systematically evaluate the MIR task performances with 13 different MIR tasks. Our findings suggest that training with music data can generally improve performance on MIR tasks, even when models are trained using paradigms designed for speech. However, we identify the limitations of such existing speech-oriented designs, especially in modelling polyphonic information. Based on the experimental results, empirical suggestions are also given for designing future musical SSL strategies and paradigms.


Vacaspati: A Diverse Corpus of Bangla Literature

arXiv.org Artificial Intelligence

Bangla (or Bengali) is the fifth most spoken language globally; yet, the state-of-the-art NLP in Bangla is lagging for even simple tasks such as lemmatization, POS tagging, etc. This is partly due to lack of a varied quality corpus. To alleviate this need, we build Vacaspati, a diverse corpus of Bangla literature. The literary works are collected from various websites; only those works that are publicly available without copyright violations or restrictions are collected. We believe that published literature captures the features of a language much better than newspapers, blogs or social media posts which tend to follow only a certain literary pattern and, therefore, miss out on language variety. Our corpus Vacaspati is varied from multiple aspects, including type of composition, topic, author, time, space, etc. It contains more than 11 million sentences and 115 million words. We also built a word embedding model, Vac-FT, using FastText from Vacaspati as well as trained an Electra model, Vac-BERT, using the corpus. Vac-BERT has far fewer parameters and requires only a fraction of resources compared to other state-of-the-art transformer models and yet performs either better or similar on various downstream tasks. On multiple downstream tasks, Vac-FT outperforms other FastText-based models. We also demonstrate the efficacy of Vacaspati as a corpus by showing that similar models built from other corpora are not as effective. The models are available at https://bangla.iitk.ac.in/.


The best Amazon device deals for Prime Day 2023

Daily Mail - Science & tech

Products featured in this Mail Best article are selected by our shopping writers. If you make a purchase using links on this page, Dailymail.co.uk will earn an affiliate commission. Amazon Prime Day has finally arrived, and there are incredible deals on tons of devices, including Ring Battery Video Doorbell Plus and the Fire HD 10 Tablet, which are at their lowest prices ever - it's a sale you won't want to pass up. This exclusive membership will grant you unrestricted access to a plethora of Prime Day deals, alongside a host of other enticing perks, including complimentary two-day shipping. You can easily cancel at any time during the 30-day trial period without any charges!


How AI could help local newsrooms remain afloat in a sea of misinformation

Engadget

It didn't take long for the downsides of a generative AI-empowered newsroom to make themselves obvious, between CNet's secret chatbot reviews editor last November and Buzzfeed's subsequent mass layoffs of human staff in favor of AI-generated "content" creators. The specter of being replaced by a "good enough AI" looms large in many a journalist's mind these days with as many as a third of the nation's newsrooms expected to shutter by the middle of the decade. But AI doesn't have to necessarily be an existential threat to the field. As six research teams showed at NYU Media Lab's AI & Local News Initiative demo day in late June, the technology may also be the key to foundationally transforming the way local news is gathered and produced. Now in its second year, the initiative is tasked with helping local news organizations to "harness the power of artificial intelligence to drive success." It's backed as part of a larger $3 million grant from the Knight Foundation which is funding four such programs in total in partnership with the Associated Press, Brown Institute's Local News Lab, NYC Media Lab and the Partnership on AI.


'Alarming' misuse of AI to spy on activists, journalists 'under guise of preventing terrorism': UN expert

FOX News

AGI, while powerful, could have negative consequences, warned Diveplane CEO Mike Capps and Liberty Blockchain CCO Christopher Alexander. A United Nations expert warned about an "alarming" trend of "using security rhetoric" to justify "intrusive and high-risk technologies," including artificial intelligence, to spy on social rights activists and journalists. U.N. expert Fionnuala Ní Aoláin called for a moratorium on AI development, among other advanced technologies like drones, until "adequate safeguards are in place," according to a March 2023 report that was presented to the Human Rights Council. "Exceptional justifications for the use of surveillance technologies in human rights'lite' counter-terrorism often turn into mundane regular use," Ní Aoláin said in a statement after the report's release. Without meaningful oversight, she argued, countries and private actors can use AI-power tech with impunity "under the guise of preventing terrorism." Fionnuala Ní Aoláin called for a moratorium on AI development, among other advanced technologies, until "adequate safeguards are in place."