Oceania
Underspecification in Scene Description-to-Depiction Tasks
Hutchinson, Ben, Baldridge, Jason, Prabhakaran, Vinodkumar
Questions regarding implicitness, ambiguity and underspecification are crucial for understanding the task validity and ethical concerns of multimodal image+text systems, yet have received little attention to date. This position paper maps out a conceptual framework to address this gap, focusing on systems which generate images depicting scenes from scene descriptions. In doing so, we account for how texts and images convey meaning differently. We outline a set of core challenges concerning textual and visual ambiguity, as well as risks that may be amplified by ambiguous and underspecified elements. We propose and discuss strategies for addressing these challenges, including generating visually ambiguous images, and generating a set of diverse images.
Short-term prediction of stream turbidity using surrogate data and a meta-model approach
Rele, Bhargav, Hogan, Caleb, Kandanaarachchi, Sevvandi, Leigh, Catherine
Many water-quality monitoring programs aim to measure turbidity to help guide effective management of waterways and catchments, yet distributing turbidity sensors throughout networks is typically cost prohibitive. To this end, we built and compared the ability of dynamic regression (ARIMA), long short-term memory neural nets (LSTM), and generalized additive models (GAM) to forecast stream turbidity one step ahead, using surrogate data from relatively low-cost in-situ sensors and publicly available databases. We iteratively trialled combinations of four surrogate covariates (rainfall, water level, air temperature and total global solar exposure) selecting a final model for each type that minimised the corrected Akaike Information Criterion. Cross-validation using a rolling time-window indicated that ARIMA, which included the rainfall and water-level covariates only, produced the most accurate predictions, followed closely by GAM, which included all four covariates. We constructed a meta-model, trained on time-series features of turbidity, to take advantage of the strengths of each model over different time points and predict the best model (that with the lowest forecast error one-step prior) for each time step. The meta-model outperformed all other models, indicating that this methodology can yield high accuracy and may be a viable alternative to using measurements sourced directly from turbidity-sensors where costs prohibit their deployment and maintenance, and when predicting turbidity across the short term. Our findings also indicated that temperature and light-associated variables, for example underwater illuminance, may hold promise as cost-effective, high-frequency surrogates of turbidity, especially when combined with other covariates, like rainfall, that are typically measured at coarse levels of spatial resolution.
Adversarial Robustness of Deep Neural Networks: A Survey from a Formal Verification Perspective
Meng, Mark Huasong, Bai, Guangdong, Teo, Sin Gee, Hou, Zhe, Xiao, Yan, Lin, Yun, Dong, Jin Song
Neural networks have been widely applied in security applications such as spam and phishing detection, intrusion prevention, and malware detection. This black-box method, however, often has uncertainty and poor explainability in applications. Furthermore, neural networks themselves are often vulnerable to adversarial attacks. For those reasons, there is a high demand for trustworthy and rigorous methods to verify the robustness of neural network models. Adversarial robustness, which concerns the reliability of a neural network when dealing with maliciously manipulated inputs, is one of the hottest topics in security and machine learning. In this work, we survey existing literature in adversarial robustness verification for neural networks and collect 39 diversified research works across machine learning, security, and software engineering domains. We systematically analyze their approaches, including how robustness is formulated, what verification techniques are used, and the strengths and limitations of each technique. We provide a taxonomy from a formal verification perspective for a comprehensive understanding of this topic. We classify the existing techniques based on property specification, problem reduction, and reasoning strategies. We also demonstrate representative techniques that have been applied in existing studies with a sample model. Finally, we discuss open questions for future research.
WinoGAViL: Gamified Association Benchmark to Challenge Vision-and-Language Models
Bitton, Yonatan, Guetta, Nitzan Bitton, Yosef, Ron, Elovici, Yuval, Bansal, Mohit, Stanovsky, Gabriel, Schwartz, Roy
While vision-and-language models perform well on tasks such as visual question answering, they struggle when it comes to basic human commonsense reasoning skills. In this work, we introduce WinoGAViL: an online game of vision-and-language associations (e.g., between werewolves and a full moon), used as a dynamic evaluation benchmark. Inspired by the popular card game Codenames, a spymaster gives a textual cue related to several visual candidates, and another player tries to identify them. Human players are rewarded for creating associations that are challenging for a rival AI model but still solvable by other human players. We use the game to collect 3.5K instances, finding that they are intuitive for humans (>90% Jaccard index) but challenging for state-of-the-art AI models, where the best model (ViLT) achieves a score of 52%, succeeding mostly where the cue is visually salient. Our analysis as well as the feedback we collect from players indicate that the collected associations require diverse reasoning skills, including general knowledge, common sense, abstraction, and more. We release the dataset, the code and the interactive game, allowing future data collection that can be used to develop models with better association abilities.
A New Look and Convergence Rate of Federated Multi-Task Learning with Laplacian Regularization
Dinh, Canh T., Vu, Tung T., Tran, Nguyen H., Dao, Minh N., Zhang, Hongyu
Non-Independent and Identically Distributed (non- IID) data distribution among clients is considered as the key factor that degrades the performance of federated learning (FL). Several approaches to handle non-IID data such as personalized FL and federated multi-task learning (FMTL) are of great interest to research communities. In this work, first, we formulate the FMTL problem using Laplacian regularization to explicitly leverage the relationships among the models of clients for multi-task learning. Then, we introduce a new view of the FMTL problem, which in the first time shows that the formulated FMTL problem can be used for conventional FL and personalized FL. We also propose two algorithms FedU and dFedU to solve the formulated FMTL problem in communication-centralized and decentralized schemes, respectively. Theoretically, we prove that the convergence rates of both algorithms achieve linear speedup for strongly convex and sublinear speedup of order 1/2 for nonconvex objectives. Experimentally, we show that our algorithms outperform the algorithm FedAvg, FedProx, SCAFFOLD, and AFL in FL settings, MOCHA in FMTL settings, as well as pFedMe and Per-FedAvg in personalized FL settings.
Girls with brothers are no more likely to grow up as 'tomboys', study finds
Scientists have rubbished the theory that a child's personality is influenced by the gender of their siblings. Many think that children who grow up around multiple siblings of the opposite sex are influenced by them in terms of personality well into adulthood. For example, girls with brothers are seen as likely to become'tomboys', while boys with sisters are seen as likely to become'girlish', according to common belief. But a new study suggests this way of thinking is a misconception – and that sibling gender'does not systematically affect personality'. Our personality as adults is not determined by whether we grow up with sisters or brothers, the new study claims.
Credit Clear share price jumps on insurer demand - Insurtech - Insurance News - insuranceNEWS.com.au
Shares in ASX-listed Credit Clear spiked 18% last week after it revealed it signed four contracts with car insurance clients. The insurtech entered new agreements with Zurich, Aioi Nissay Dowa and another motor insurance specialist last month, and expanded an existing relationship with a fourth large insurance group. It expects to announce more insurance clients in the coming months. Credit Clear – a finalist in the 2022 Australian and New Zealand Institute of Insurance and Finance (ANZIIF) industry awards – offers products based on AI models, automation and predictive analytics. It recently developed an at-fault third party claim system for car insurers in collaboration with a large Australian insurer.
Samsung's Tizen OS is coming to other brands' TVs
Last week LG announced that it would allow third-party TV manufacturers to use its webOS platform and now its main rival is following suit. Samsung has revealed that it will license its Tizen OS TV platform for use in non-Samsung TV models for the first time, partnering with Akai, RCA and a bunch of other brands (Bauhn, Linsar, Sunny, Vispera) sold in Europe, Australia and New Zealand. The partnership gives those manufacturers access to Tizen OS features like Samsung TV Plus (a free streaming TV and video platform), Universal Guide for discovery and personalized recommendations, and Samsung's Bixby and other voice assistants. As we noted when LG first announced it would license webOS to other TV makers, these deals give buyers another option on lower-priced smart TVs that might otherwise run Android TV, Roku or Amazon's Fire TV. While you've probably never heard of many of the brands mentioned, the fact that Samsung is opening its Tizen platform means it could come to TVs sold in the US at some point.
Dynamic Dialogue Policy for Continual Reinforcement Learning
Geishauser, Christian, van Niekerk, Carel, Lubis, Nurul, Heck, Michael, Lin, Hsien-Chin, Feng, Shutong, Gašić, Milica
Continual learning is one of the key components of human learning and a necessary requirement of artificial intelligence. As dialogue can potentially span infinitely many topics and tasks, a task-oriented dialogue system must have the capability to continually learn, dynamically adapting to new challenges while preserving the knowledge it already acquired. Despite the importance, continual reinforcement learning of the dialogue policy has remained largely unaddressed. The lack of a framework with training protocols, baseline models and suitable metrics, has so far hindered research in this direction. In this work we fill precisely this gap, enabling research in dialogue policy optimisation to go from static to dynamic learning. We provide a continual learning algorithm, baseline architectures and metrics for assessing continual learning models. Moreover, we propose the dynamic dialogue policy transformer (DDPT), a novel dynamic architecture that can integrate new knowledge seamlessly, is capable of handling large state spaces and obtains significant zero-shot performance when being exposed to unseen domains, without any growth in network parameter size.
Checks and Strategies for Enabling Code-Switched Machine Translation
Gowda, Thamme, Gheini, Mozhdeh, May, Jonathan
Code-switching is a common phenomenon among multilingual speakers, where alternation between two or more languages occurs within the context of a single conversation. While multilingual humans can seamlessly switch back and forth between languages, multilingual neural machine translation (NMT) models are not robust to such sudden changes in input. This work explores multilingual NMT models' ability to handle code-switched text. First, we propose checks to measure switching capability. Second, we investigate simple and effective data augmentation methods that can enhance an NMT model's ability to support code-switching. Finally, by using a glass-box analysis of attention modules, we demonstrate the effectiveness of these methods in improving robustness.