Africa
Do LLMs Work on Charts? Designing Few-Shot Prompts for Chart Question Answering and Summarization
Do, Xuan Long, Hassanpour, Mohammad, Masry, Ahmed, Kavehzadeh, Parsa, Hoque, Enamul, Joty, Shafiq
A number of tasks have been proposed recently to facilitate easy access to charts such as chart QA and summarization. The dominant paradigm to solve these tasks has been to fine-tune a pretrained model on the task data. However, this approach is not only expensive but also not generalizable to unseen tasks. On the other hand, large language models (LLMs) have shown impressive generalization capabilities to unseen tasks with zero- or few-shot prompting. However, their application to chart-related tasks is not trivial as these tasks typically involve considering not only the underlying data but also the visual features in the chart image. We propose PromptChart, a multimodal few-shot prompting framework with LLMs for chart-related applications. By analyzing the tasks carefully, we have come up with a set of prompting guidelines for each task to elicit the best few-shot performance from LLMs. We further propose a strategy to inject visual information into the prompts. Our experiments on three different chart-related information consumption tasks show that with properly designed prompts LLMs can excel on the benchmarks, achieving state-of-the-art.
RTQ: Rethinking Video-language Understanding Based on Image-text Model
Wang, Xiao, Li, Yaoyu, Gan, Tian, Zhang, Zheng, Lv, Jingjing, Nie, Liqiang
Recent advancements in video-language understanding have been established on the foundation of image-text models, resulting in promising outcomes due to the shared knowledge between images and videos. However, video-language understanding presents unique challenges due to the inclusion of highly complex semantic details, which result in information redundancy, temporal dependency, and scene complexity. Current techniques have only partially tackled these issues, and our quantitative analysis indicates that some of these methods are complementary. In light of this, we propose a novel framework called RTQ (Refine, Temporal model, and Query), which addresses these challenges simultaneously. The approach involves refining redundant information within frames, modeling temporal relations among frames, and querying task-specific information from the videos. Remarkably, our model demonstrates outstanding performance even in the absence of video-language pre-training, and the results are comparable with or superior to those achieved by state-of-the-art pre-training methods. Code is available at https://github.com/SCZwangxiao/RTQ-MM2023.
PIGEON: Predicting Image Geolocations
Haas, Lukas, Skreta, Michal, Alberti, Silas, Finn, Chelsea
Planet-scale image geolocalization remains a challenging problem due to the diversity of images originating from anywhere in the world. Although approaches based on vision transformers have made significant progress in geolocalization accuracy, success in prior literature is constrained to narrow distributions of images of landmarks, and performance has not generalized to unseen places. We present a new geolocalization system that combines semantic geocell creation, multi-task contrastive pretraining, and a novel loss function. Additionally, our work is the first to perform retrieval over location clusters for guess refinements. We train two models for evaluations on street-level data and general-purpose image geolocalization; the first model, PIGEON, is trained on data from the game of Geoguessr and is capable of placing over 40% of its guesses within 25 kilometers of the target location globally. We also develop a bot and deploy PIGEON in a blind experiment against humans, ranking in the top 0.01% of players. We further challenge one of the world's foremost professional Geoguessr players to a series of six matches with millions of viewers, winning all six games. Our second model, PIGEOTTO, differs in that it is trained on a dataset of images from Flickr and Wikipedia, achieving state-of-the-art results on a wide range of image geolocalization benchmarks, outperforming the previous SOTA by up to 7.7 percentage points on the city accuracy level and up to 38.8 percentage points on the country level. Our findings suggest that PIGEOTTO is the first image geolocalization model that effectively generalizes to unseen places and that our approach can pave the way for highly accurate, planet-scale image geolocalization systems. Our code is available on GitHub.
T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Mixed Large Language Model Signals for Science Question Answering
Wang, Lei, Hu, Yi, He, Jiabang, Xu, Xing, Liu, Ning, Liu, Hui, Shen, Heng Tao
Large Language Models (LLMs) have recently demonstrated exceptional performance in various Natural Language Processing (NLP) tasks. They have also shown the ability to perform chain-of-thought (CoT) reasoning to solve complex problems. Recent studies have explored CoT reasoning in complex multimodal scenarios, such as the science question answering task, by fine-tuning multimodal models with high-quality human-annotated CoT rationales. However, collecting high-quality COT rationales is usually time-consuming and costly. Besides, the annotated rationales are hardly accurate due to the external essential information missed. To address these issues, we propose a novel method termed T-SciQ that aims at teaching science question answering with LLM signals. The T-SciQ approach generates high-quality CoT rationales as teaching signals and is advanced to train much smaller models to perform CoT reasoning in complex modalities. Additionally, we introduce a novel data mixing strategy to produce more effective teaching data samples for simple and complex science question answer problems. Extensive experimental results show that our T-SciQ method achieves a new state-of-the-art performance on the ScienceQA benchmark, with an accuracy of 96.18%. Moreover, our approach outperforms the most powerful fine-tuned baseline by 4.5%. The code is publicly available at https://github.com/T-SciQ/T-SciQ.
US, UK say they shot down 15 drones from Yemen's Houthis over Red Sea
The United States and United Kingdom authorities say their warships have shot down 15 attack drones over the Red Sea as Israel's war on Gaza threatens to spread in the region. The US Central Command (CENTCOM) on Saturday said its guided-missile destroyer responded to a wave of drones from "Houthi-controlled areas of Yemen" over the Red Sea, downing 14 suspected attack drones. It described the launches as "one-way attack drones", saying they were "shot down with no damage to ships in the area or reported injuries". UK Defence Secretary Grant Shapps also said the Royal Navy destroyer HMS Diamond fired a Sea Viper missile and destroyed a drone that was "targeting merchant shipping". Meanwhile, Yemen's Iran-aligned Houthis said the group attacked the Israeli city of Eilat on Saturday with a swarm of drones, according to spokesman Yahya Sarea who referred to the Red Sea resort city as being in "southern occupied Palestine".
Physicist Bob Coecke: 'It's easier to convince kids than adults about quantum mechanics'
Belgian physicist and musician Prof Bob Coecke, 55, wants to teach quantum physics to a mass audience. The paradox-filled theory that describes the microscopic realm has become a staple of science fiction, from Marvel's Ant-Man to the multiple Oscar-winning Everything Everywhere All at Once. It's famously bizarre and, in the UK, the subject is mostly reserved for undergraduates specialising in physics because it requires grappling with complicated maths. But Coecke, a former Oxford professor, has devised a maths-free framework using diagrams for total beginners, outlined in Quantum in Pictures, his book with Dr Stefano Gogioso that was published earlier this year. Over the summer, they ran an education experiment, teaching the pictorial method to UK schoolchildren – who then beat the average exam scores of Oxford University's postgraduate physics students.
Will oil prices rise after Red Sea shipping curbs amid Houthi attacks?
Hijackings, missile strikes and drone assaults on ships by Yemen's Houthi rebels have forced AP Moller-Maersk, a Danish shipping and logistics giant, and Hapag-Lloyd, a German shipping and container transportation company, to pause shipments through the Red Sea. Their decisions, announced on Friday, are a sign that major corporations are taking the security situation in the Red Sea increasingly seriously. But the consequences might also be felt by the world's oil markets and the cost of energy that consumers need to bear – though the extent of any disruption might depend on how major global players respond to the looming crisis, said experts. Maersk said in a statement that its decision stemmed from the company's concerns about the "highly escalated security situation in the southern Red Sea and Gulf of Aden" over the past few weeks. Recent missile and drone attacks on commercial vessels represent a "significant threat to the safety and security of seafarers," it said.
A new method color MS-BSIF Features learning for the robust kinship verification
Aliradi, Rachid, Ouamane, Abdealmalik, Amrane, Abdeslam
the paper presents a new method color MS-BSIF learning and MS-LBP for the kinship verification is the machine's ability to identify the genetic and blood the relationship and its degree between the facial images of humans. Facial verification of kinship refers to the task of training a machine to recognize the blood relationship between a pair of faces parent and non-parent (verification) based on features extracted from facial images, and determining the exact type or degree of this genetic relationship. We use the LBP and color BSIF learning features for the comparison and the TXQDA method for dimensionality reduction and data classification. We let's test the kinship facial verification application is namely the kinface Cornell database. This system improves the robustness of learning while controlling efficiency. The experimental results obtained and compared to other methods have proven the reliability of our framework and surpass the performance of other state-of-the-art techniques.
Sentiment Analysis and Text Analysis of the Public Discourse on Twitter about COVID-19 and MPox
Mining and analysis of the big data of Twitter conversations have been of significant interest to the scientific community in the fields of healthcare, epidemiology, big data, data science, computer science, and their related areas, as can be seen from several works in the last few years that focused on sentiment analysis and other forms of text analysis of tweets related to Ebola, E-Coli, Dengue, Human Papillomavirus, Middle East Respiratory Syndrome, Measles, Zika virus, H1N1, influenza like illness, swine flu, flu, Cholera, Listeriosis, cancer, Liver Disease, Inflammatory Bowel Disease, kidney disease, lupus, Parkinsons, Diphtheria, and West Nile virus. The recent outbreaks of COVID-19 and MPox have served as catalysts for Twitter usage related to seeking and sharing information, views, opinions, and sentiments involving both of these viruses. None of the prior works in this field analyzed tweets focusing on both COVID-19 and MPox simultaneously. To address this research gap, a total of 61,862 tweets that focused on MPox and COVID-19 simultaneously, posted between 7 May 2022 and 3 March 2023, were studied. The findings and contributions of this study are manifold. First, the results of sentiment analysis using the VADER approach show that nearly half the tweets had a negative sentiment. It was followed by tweets that had a positive sentiment and tweets that had a neutral sentiment, respectively. Second, this paper presents the top 50 hashtags used in these tweets. Third, it presents the top 100 most frequently used words in these tweets after performing tokenization, removal of stopwords, and word frequency analysis. Finally, a comprehensive comparative study that compares the contributions of this paper with 49 prior works in this field is presented to further uphold the relevance and novelty of this work.