Goto

Collaborating Authors

 Indian Ocean


ViWikiFC: Fact-Checking for Vietnamese Wikipedia-Based Textual Knowledge Source

arXiv.org Artificial Intelligence

Fact-checking is essential due to the explosion of misinformation in the media ecosystem. Although false information exists in every language and country, most research to solve the problem mainly concentrated on huge communities like English and Chinese. Low-resource languages like Vietnamese are necessary to explore corpora and models for fact verification. To bridge this gap, we construct ViWikiFC, the first manual annotated open-domain corpus for Vietnamese Wikipedia Fact Checking more than 20K claims generated by converting evidence sentences extracted from Wikipedia articles. We analyze our corpus through many linguistic aspects, from the new dependency rate, the new n-gram rate, and the new word rate. We conducted various experiments for Vietnamese fact-checking, including evidence retrieval and verdict prediction. BM25 and InfoXLM (Large) achieved the best results in two tasks, with BM25 achieving an accuracy of 88.30% for SUPPORTS, 86.93% for REFUTES, and only 56.67% for the NEI label in the evidence retrieval task, InfoXLM (Large) achieved an F1 score of 86.51%. Furthermore, we also conducted a pipeline approach, which only achieved a strict accuracy of 67.00% when using InfoXLM (Large) and BM25. These results demonstrate that our dataset is challenging for the Vietnamese language model in fact-checking tasks.


OXYGENERATOR: Reconstructing Global Ocean Deoxygenation Over a Century with Deep Learning

arXiv.org Artificial Intelligence

Accurately reconstructing the global ocean deoxygenation over a century is crucial for assessing and protecting marine ecosystem. Existing expert-dominated numerical simulations fail to catch up with the dynamic variation caused by global warming and human activities. Besides, due to the high-cost data collection, the historical observations are severely sparse, leading to big challenge for precise reconstruction. In this work, we propose OxyGenerator, the first deep learning based model, to reconstruct the global ocean deoxygenation from 1920 to 2023. Specifically, to address the heterogeneity across large temporal and spatial scales, we propose zoning-varying graph message-passing to capture the complex oceanographic correlations between missing values and sparse observations. Additionally, to further calibrate the uncertainty, we incorporate inductive bias from dissolved oxygen (DO) variations and chemical effects. Compared with in-situ DO observations, OxyGenerator significantly outperforms CMIP6 numerical simulations, reducing MAPE by 38.77%, demonstrating a promising potential to understand the "breathless ocean" in data-driven manner.


SaudiBERT: A Large Language Model Pretrained on Saudi Dialect Corpora

arXiv.org Artificial Intelligence

In this paper, we introduce SaudiBERT, a monodialect Arabic language model pretrained exclusively on Saudi dialectal text. To demonstrate the model's effectiveness, we compared SaudiBERT with six different multidialect Arabic language models across 11 evaluation datasets, which are divided into two groups: sentiment analysis and text classification. SaudiBERT achieved average F1-scores of 86.15\% and 87.86\% in these groups respectively, significantly outperforming all other comparative models. Additionally, we present two novel Saudi dialectal corpora: the Saudi Tweets Mega Corpus (STMC), which contains over 141 million tweets in Saudi dialect, and the Saudi Forums Corpus (SFC), which includes 15.2 GB of text collected from five Saudi online forums. Both corpora are used in pretraining the proposed model, and they are the largest Saudi dialectal corpora ever reported in the literature. The results confirm the effectiveness of SaudiBERT in understanding and analyzing Arabic text expressed in Saudi dialect, achieving state-of-the-art results in most tasks and surpassing other language models included in the study. SaudiBERT model is publicly available on \url{https://huggingface.co/faisalq/SaudiBERT}.


China launches lunar probe to take samples from far side of the moon

FOX News

Former National Security Adviser Robert O'Brien joins'Life, Liberty & Levin' to discuss the Biden administration's foreign policy in the Middle East. China on Friday launched a lunar probe to land on the far side of the moon and return with samples that could provide insights into differences between the less-explored region and the better-known near side. It is the latest advance in China's increasingly sophisticated space exploration program, which is now competing with the U.S., still the leader in space. China also has a three-member crew on its own orbiting space station and aims to put astronauts on the moon by 2030. Three Chinese lunar probe missions are planned over the next four years.


Portuguese-flagged ship targeted in Arabian Sea drone assault; Houthi rebels claim responsibility

FOX News

Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. A Portuguese-flagged container ship came under attack by a drone in the far reaches of the Arabian Sea, corresponding with a claim by Yemen's Houthi rebels that they assaulted the ship there, authorities said Tuesday. The attack on the MSC Orion, occurring some 375 miles off the coast of Yemen, appeared to be the first confirmed deep-sea assault claimed by the Houthis since they began targeting ships in November. It suggests the Houthis -- or potentially their main benefactor Iran -- may have the ability to strike into the distances of the Indian Ocean as the rebels previously threatened in their ongoing campaign over Israel's war on Hamas in the Gaza Strip.


Likely missile attack by Yemen's Houthi rebels damages a ship in the Red Sea

FOX News

Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. A suspected missile attack by Yemen's Houthi rebels damaged a ship in the Red Sea on Monday, authorities said, the latest assault in their campaign against international shipping in the crucial maritime route. The attack happened off the coast of Mokha, Yemen, the British military's United Kingdom Maritime Trade Operations center said. The ship sustained damage in the attack, the UKMTO said, though its crew was safe and heading to its next port of call.


Automated Construction of Theme-specific Knowledge Graphs

arXiv.org Artificial Intelligence

Despite widespread applications of knowledge graphs (KGs) in various tasks such as question answering and intelligent conversational systems, existing KGs face two major challenges: information granularity and deficiency in timeliness. These hinder considerably the retrieval and analysis of in-context, fine-grained, and up-to-date knowledge from KGs, particularly in highly specialized themes (e.g., specialized scientific research) and rapidly evolving contexts (e.g., breaking news or disaster tracking). To tackle such challenges, we propose a theme-specific knowledge graph (i.e., ThemeKG), a KG constructed from a theme-specific corpus, and design an unsupervised framework for ThemeKG construction (named TKGCon). The framework takes raw theme-specific corpus and generates a high-quality KG that includes salient entities and relations under the theme. Specifically, we start with an entity ontology of the theme from Wikipedia, based on which we then generate candidate relations by Large Language Models (LLMs) to construct a relation ontology. To parse the documents from the theme corpus, we first map the extracted entity pairs to the ontology and retrieve the candidate relations. Finally, we incorporate the context and ontology to consolidate the relations for entity pairs. We observe that directly prompting GPT-4 for theme-specific KG leads to inaccurate entities (such as "two main types" as one entity in the query result) and unclear (such as "is", "has") or wrong relations (such as "have due to", "to start"). In contrast, by constructing the theme-specific KG step by step, our model outperforms GPT-4 and could consistently identify accurate entities and relations. Experimental results also show that our framework excels in evaluations compared with various KG construction baselines.


Iranian-backed Houthis claim responsibility for US reaper drone crash off Yemen coast

FOX News

Iranian-backed Houthis rebels have claimed responsibility for a U.S. MQ-9 Reaper drone crash off the coast of Yemen on Thursday, Fox News confirmed on Friday. Thursday's crash is the fourth remotely piloted drone brought down by Iranian-proxy groups since November, costing the U.S. government upwards of 120 million. It is also the third time Houthi rebels have brought down a U.S. MQ-9 drone. Remotely piloted MQ-9 Reaper drones cost around 30 million each. Last fall, the Houthis released video of a reaper drone the rebels shot down on Nov. 8, one day after Hamas' unprovoked attack on Israel.


TIGQA:An Expert Annotated Question Answering Dataset in Tigrinya

arXiv.org Artificial Intelligence

The absence of explicitly tailored, accessible annotated datasets for educational purposes presents a notable obstacle for NLP tasks in languages with limited resources.This study initially explores the feasibility of using machine translation (MT) to convert an existing dataset into a Tigrinya dataset in SQuAD format. As a result, we present TIGQA, an expert annotated educational dataset consisting of 2.68K question-answer pairs covering 122 diverse topics such as climate, water, and traffic. These pairs are from 537 context paragraphs in publicly accessible Tigrinya and Biology books. Through comprehensive analyses, we demonstrate that the TIGQA dataset requires skills beyond simple word matching, requiring both single-sentence and multiple-sentence inference abilities. We conduct experiments using state-of-the art MRC methods, marking the first exploration of such models on TIGQA. Additionally, we estimate human performance on the dataset and juxtapose it with the results obtained from pretrained models.The notable disparities between human performance and best model performance underscore the potential for further enhancements to TIGQA through continued research. Our dataset is freely accessible via the provided link to encourage the research community to address the challenges in the Tigrinya MRC.


Tesla profit plunges 55%, as shares bounce on plans for cheaper vehicles

Al Jazeera

Tesla reported a 55 percent drop in profit amid fierce competition in the electric vehicle market, but shares rallied on plans to accelerate the production of more affordable models. The Austin, Texas-based company on Tuesday reported profits of 1.1bn in the first quarter, down from 2.51bn a year ago. But shares of Tesla soared by 11 percent after CEO Elon Musk said that production of new, more affordable vehicles would begin in the second half of next year "if not late this year". The models "will use new aspects of the next generation platform as well as aspects of our current platform", Musk said on a conference call with analysts. Musk did not elaborate on the new vehicles, saying more details would be released in August.