Oceania
Accelerated Riemannian Optimization: Handling Constraints with a Prox to Bound Geometric Penalties
Martínez-Rubio, David, Pokutta, Sebastian
We propose a globally-accelerated, first-order method for the optimization of smooth and (strongly or not) geodesically-convex functions in a wide class of Hadamard manifolds. We achieve the same convergence rates as Nesterov's accelerated gradient descent, up to a multiplicative geometric penalty and log factors. Crucially, we can enforce our method to stay within a compact set we define. Prior fully accelerated works \emph{resort to assuming} that the iterates of their algorithms stay in some pre-specified compact set, except for two previous methods of limited applicability. For our manifolds, this solves the open question in [KY22] about obtaining global general acceleration without iterates assumptively staying in the feasible set. In our solution, we design an accelerated Riemannian inexact proximal point algorithm, which is a result that was unknown even with exact access to the proximal operator, and is of independent interest. For smooth functions, we show we can implement the prox step inexactly with first-order methods in Riemannian balls of certain diameter that is enough for global accelerated optimization.
A partial order view of message-passing communication models
Di Giusto, Cinzia, Ferré, Davide, Laversa, Laetitia, Lozes, Etienne
There is a wide variety of message-passing communication models, ranging from synchronous ''rendez-vous'' communications to fully asynchronous/out-of-order communications. For large-scale distributed systems, the communication model is determined by the transport layer of the network, and a few classes of orders of message delivery (FIFO, causally ordered) have been identified in the early days of distributed computing. For local-scale message-passing applications, e.g., running on a single machine, the communication model may be determined by the actual implementation of message buffers and by how FIFO queues are used. While large-scale communication models, such as causal ordering, are defined by logical axioms, local-scale models are often defined by an operational semantics. In this work, we connect these two approaches, and we present a unified hierarchy of communication models encompassing both large-scale and local-scale models, based on their concurrent behaviors. We also show that all the communication models we consider can be axiomatized in the monadic second order logic, and may therefore benefit from several bounded verification techniques based on bounded special treewidth.
The road to autonomous vehicles is via industrial usage – and lots of data
For me, autonomous vehicles inspire profound fascination and deep fear. I'm fascinated by their prospects for the future--their inevitability. I also constantly wonder: how can we make this mode of travel safe? How's it possible a self-driving car could ever be safe? At CES 2023, I started to get some answers.
Extracting Medication Changes in Clinical Narratives using Pre-trained Language Models
Ramachandran, Giridhar Kaushik, Lybarger, Kevin, Liu, Yaya, Mahajan, Diwakar, Liang, Jennifer J., Tsou, Ching-Huei, Yetisgen, Meliha, Uzuner, Özlem
An accurate and detailed account of patient medications, including medication changes within the patient timeline, is essential for healthcare providers to provide appropriate patient care. Healthcare providers or the patients themselves may initiate changes to patient medication. Medication changes take many forms, including prescribed medication and associated dosage modification. These changes provide information about the overall health of the patient and the rationale that led to the current care. Future care can then build on the resulting state of the patient. This work explores the automatic extraction of medication change information from free-text clinical notes. The Contextual Medication Event Dataset (CMED) is a corpus of clinical notes with annotations that characterize medication changes through multiple change-related attributes, including the type of change (start, stop, increase, etc.), initiator of the change, temporality, change likelihood, and negation. Using CMED, we identify medication mentions in clinical text and propose three novel high-performing BERT-based systems that resolve the annotated medication change characteristics. We demonstrate that our proposed systems improve medication change classification performance over the initial work exploring CMED.
SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
Tanaka, Ryota, Nishida, Kyosuke, Nishida, Kosuke, Hasegawa, Taku, Saito, Itsumi, Saito, Kuniko
Visual question answering on document images that contain textual, visual, and layout information, called document VQA, has received much attention recently. Although many datasets have been proposed for developing document VQA systems, most of the existing datasets focus on understanding the content relationships within a single image and not across multiple images. In this study, we propose a new multi-image document VQA dataset, SlideVQA, containing 2.6k+ slide decks composed of 52k+ slide images and 14.5k questions about a slide deck. SlideVQA requires complex reasoning, including single-hop, multi-hop, and numerical reasoning, and also provides annotated arithmetic expressions of numerical answers for enhancing the ability of numerical reasoning. Moreover, we developed a new end-to-end document VQA model that treats evidence selection and question answering in a unified sequence-to-sequence format. Experiments on SlideVQA show that our model outperformed existing state-of-the-art QA models, but that it still has a large gap behind human performance. We believe that our dataset will facilitate research on document VQA.
TransfQMix: Transformers for Leveraging the Graph Structure of Multi-Agent Reinforcement Learning Problems
Gallici, Matteo, Martin, Mario, Masmitja, Ivan
Coordination is one of the most difficult aspects of multi-agent reinforcement learning (MARL). One reason is that agents normally choose their actions independently of one another. In order to see coordination strategies emerging from the combination of independent policies, the recent research has focused on the use of a centralized function (CF) that learns each agent's contribution to the team reward. However, the structure in which the environment is presented to the agents and to the CF is typically overlooked. We have observed that the features used to describe the coordination problem can be represented as vertex features of a latent graph structure. Here, we present TransfQMix, a new approach that uses transformers to leverage this latent structure and learn better coordination policies. Our transformer agents perform a graph reasoning over the state of the observable entities. Our transformer Q-mixer learns a monotonic mixing-function from a larger graph that includes the internal and external states of the agents. TransfQMix is designed to be entirely transferable, meaning that same parameters can be used to control and train larger or smaller teams of agents. This enables to deploy promising approaches to save training time and derive general policies in MARL, such as transfer learning, zero-shot transfer, and curriculum learning. We report TransfQMix's performances in the Spread and StarCraft II environments. In both settings, it outperforms state-of-the-art Q-Learning models, and it demonstrates effectiveness in solving problems that other methods can not solve.
Experimental System Identification and Disturbance Observer-based Control for a Monolithic $Z{\theta}_{x}{\theta}_{y}$ Precision Positioning System
Ghafarian, Mohammadali, Shirinzadeh, Bijan, Al-Jodah, Ammar, Das, Tilok Kumar, Shen, Tianyao
A compliant parallel micromanipulator is a mechanism in which the moving platform is connected to the base through a number of flexural components. Utilizing parallel-kinematics configurations and flexure joints, the monolithic micromanipulators can achieve extremely high motion resolution and accuracy. In this work, the focus was towards the experimental evaluation of a 3-DOF ($Z{\theta}_{x}{\theta}_{y}$) monolithic flexure-based piezo-driven micromanipulator for precise out-of-plane micro/nano positioning applications. The monolithic structure avoids the deficiencies of non-monolithic designs such as backlash, wear, friction, and improves the performance of micromanipulator in terms of high resolution, accuracy, and repeatability. A computational study was conducted to investigate and obtain the inverse kinematics of the proposed micromanipulator. As a result of computational analysis, the developed prototype of the micromanipulator is capable of executing large motion range of $\pm$238.5$\mu$m $\times$ $\pm$4830.5$\mu$rad $\times$ $\pm$5486.2$\mu$rad. Finally, a sliding mode control strategy with nonlinear disturbance observer (SMC-NDO) was designed and implemented on the proposed micromanipulator to obtain system behaviors during experiments. The obtained results from different experimental tests validated the fine micromanipulator's positioning ability and the efficiency of the control methodology for precise micro/nano manipulation applications. The proposed micromanipulator achieved very fine spatial and rotational resolutions of $\pm$4nm, $\pm$250nrad, and $\pm$230nrad throughout its workspace.
Multiple-level Point Embedding for Solving Human Trajectory Imputation with Prediction
Qin, Kyle K., Ren, Yongli, Shao, Wei, Lake, Brennan, Privitera, Filippo, Salim, Flora D.
Sparsity is a common issue in many trajectory datasets, including human mobility data. This issue frequently brings more difficulty to relevant learning tasks, such as trajectory imputation and prediction. Nowadays, little existing work simultaneously deals with imputation and prediction on human trajectories. This work plans to explore whether the learning process of imputation and prediction could benefit from each other to achieve better outcomes. And the question will be answered by studying the coexistence patterns between missing points and observed ones in incomplete trajectories. More specifically, the proposed model develops an imputation component based on the self-attention mechanism to capture the coexistence patterns between observations and missing points among encoder-decoder layers. Meanwhile, a recurrent unit is integrated to extract the sequential embeddings from newly imputed sequences for predicting the following location. Furthermore, a new implementation called Imputation Cycle is introduced to enable gradual imputation with prediction enhancement at multiple levels, which helps to accelerate the speed of convergence. The experimental results on three different real-world mobility datasets show that the proposed approach has significant advantages over the competitive baselines across both imputation and prediction tasks in terms of accuracy and stability.
Detecting Change Intervals with Isolation Distributional Kernel
Cao, Yang, Zhu, Ye, Ting, Kai Ming, Salim, Flora D., Li, Hong Xian, Li, Gang
Detecting abrupt changes in data distribution is one of the most significant tasks in streaming data analysis. Although many unsupervised Change-Point Detection (CPD) methods have been proposed recently to identify those changes, they still suffer from missing subtle changes, poor scalability, or/and sensitive to noise points. To meet these challenges, we are the first to generalise the CPD problem as a special case of the Change-Interval Detection (CID) problem. Then we propose a CID method, named iCID, based on a recent Isolation Distributional Kernel (IDK). iCID identifies the change interval if there is a high dissimilarity score between two non-homogeneous temporal adjacent intervals. The data-dependent property and finite feature map of IDK enabled iCID to efficiently identify various types of change points in data streams with the tolerance of noise points. Moreover, the proposed online and offline versions of iCID have the ability to optimise key parameter settings. The effectiveness and efficiency of iCID have been systematically verified on both synthetic and real-world datasets.
Australian universities to return to 'pen and paper' exams after students caught using AI to write essays
Australian universities have been forced to change the way they run exams and other assessments amid fears students are using emerging artificial intelligence software to write essays. Major institutions have added new rules which state that the use of AI is cheating, with some students already caught using the software. But one AI expert has warned universities are in an "arms race" they can never win. ChatGPT, which generates text on any subject in response to a prompt or query, was launched in November by OpenAI and has already been banned across all devices in New York's public schools due to concerns over its "negative impact on student learning" and potential for plagiarism. In London, one academic tested it against a 2022 exam question and said the AI's answer was "coherent, comprehensive and sticks to the points, something students often fail to do", adding he would have to "set a different kind of exam" or deprive students of internet access for future exams. In Australia, academics have cited concerns over ChatGPT and similar technology's ability to evade anti-plagiarism software while providing quick and credible academic writing.