Goto

Collaborating Authors

 Government


Automatic Standardization of Arabic Dialects for Machine Translation

arXiv.org Artificial Intelligence

Based on an annotated multimedia corpus, television series Mar{\=a}y{\=a} 2013, we dig into the question of ''automatic standardization'' of Arabic dialects for machine translation. Here we distinguish between rule-based machine translation and statistical machine translation. Machine translation from Arabic most of the time takes standard or modern Arabic as the source language and produces quite satisfactory translations thanks to the availability of the translation memories necessary for training the models. The case is different for the translation of Arabic dialects. The productions are much less efficient. In our research we try to apply machine translation methods to a dialect/standard (or modern) Arabic pair to automatically produce a standard Arabic text from a dialect input, a process we call ''automatic standardization''. we opt here for the application of ''statistical models'' because ''automatic standardization'' based on rules is more hard with the lack of ''diglossic'' dictionaries on the one hand and the difficulty of creating linguistic rules for each dialect on the other. Carrying out this research could then lead to combining ''automatic standardization'' software and automatic translation software so that we take the output of the first software and introduce it as input into the second one to obtain at the end a quality machine translation. This approach may also have educational applications such as the development of applications to help understand different Arabic dialects by transforming dialectal texts into standard Arabic.


Physics-Informed Kernel Embeddings: Integrating Prior System Knowledge with Data-Driven Control

arXiv.org Artificial Intelligence

Data-driven control algorithms use observations of system dynamics to construct an implicit model for the purpose of control. However, in practice, data-driven techniques often require excessive sample sizes, which may be infeasible in real-world scenarios where only limited observations of the system are available. Furthermore, purely data-driven methods often neglect useful a priori knowledge, such as approximate models of the system dynamics. We present a method to incorporate such prior knowledge into data-driven control algorithms using kernel embeddings, a nonparametric machine learning technique based in the theory of reproducing kernel Hilbert spaces. Our proposed approach incorporates prior knowledge of the system dynamics as a bias term in the kernel learning problem. We formulate the biased learning problem as a least-squares problem with a regularization term that is informed by the dynamics, that has an efficiently computable, closed-form solution. Through numerical experiments, we empirically demonstrate the improved sample efficiency and out-of-sample generalization of our approach over a purely data-driven baseline. We demonstrate an application of our method to control through a target tracking problem with nonholonomic dynamics, and on spring-mass-damper and F-16 aircraft state prediction tasks.


Bayesian Additive Main Effects and Multiplicative Interaction Models using Tensor Regression for Multi-environmental Trials

arXiv.org Artificial Intelligence

We propose a Bayesian tensor regression model to accommodate the effect of multiple factors on phenotype prediction. We adopt a set of prior distributions that resolve identifiability issues that may arise between the parameters in the model. Simulation experiments show that our method out-performs previous related models and machine learning algorithms under different sample sizes and degrees of complexity. We further explore the applicability of our model by analysing real-world data related to wheat production across Ireland from 2010 to 2019. Our model performs competitively and overcomes key limitations found in other analogous approaches. Finally, we adapt a set of visualisations for the posterior distribution of the tensor effects that facilitate the identification of optimal interactions between the tensor variables whilst accounting for the uncertainty in the posterior distribution.


Automatic Differentiation of Programs with Discrete Randomness

arXiv.org Artificial Intelligence

Automatic differentiation (AD), a technique for constructing new programs which compute the derivative of an original program, has become ubiquitous throughout scientific computing and deep learning due to the improved performance afforded by gradient-based optimization. However, AD systems have been restricted to the subset of programs that have a continuous dependence on parameters. Programs that have discrete stochastic behaviors governed by distribution parameters, such as flipping a coin with probability $p$ of being heads, pose a challenge to these systems because the connection between the result (heads vs tails) and the parameters ($p$) is fundamentally discrete. In this paper we develop a new reparameterization-based methodology that allows for generating programs whose expectation is the derivative of the expectation of the original program. We showcase how this method gives an unbiased and low-variance estimator which is as automated as traditional AD mechanisms. We demonstrate unbiased forward-mode AD of discrete-time Markov chains, agent-based models such as Conway's Game of Life, and unbiased reverse-mode AD of a particle filter. Our code package is available at https://github.com/gaurav-arya/StochasticAD.jl.


Overview of Building a Model using SageMaker

#artificialintelligence

Now if you remember, next in the workflow is building the model. To go back to the hub, AKA the SageMaker dashboard, you'll notice that notebooks are next. If you're familiar with it, SageMaker notebooks are basically managed Jupyter Notebook setups. Jupyter Notebook is what IPython notebooks rebranded to a few years back, if you've never heard of it. They also compete with a service called Zeplin, but SageMaker uses a pre-installed managed version of Jupyter. Now, just know that, although you're gonna build your notebook in SageMaker, Jupyter Notebooks are actually an open-sourced application that you can download and run yourself or run your in-house servers, several SageMakers, so you're not getting locked in.



As AI rises, lawmakers try to catch up - abtlive

#artificialintelligence

From "intelligent" vacuum cleaners and driverless cars to advanced techniques for diagnosing diseases, artificial intelligence has burrowed its way into every arena of modern life. Its promoters reckon it is revolutionising human experience, but critics stress that the technology risks putting machines in charge of life-changing decisions. Regulators in Europe and North America are worried. The European Union is likely to pass legislation next year- the AI Act- aimed at reining in the age of the algorithm. The United States recently published a blueprint for an AI Bill of Rights and Canada is also mulling legislation.


Top 10 technology and ethics stories of 2022

#artificialintelligence

A major focus of Computer Weekly's technology and ethics coverage in 2022 was on working conditions throughout the tech sector, from the issue of forced labour and slavery throughout technology supply chains, to UK Amazon workers staging spontaneous "wildcat" strikes in response to derisory pay rises and warehouse conditions. Other stories in this vein included coverage of accusations that "soft union-busting" tactics were used by app-based food delivery firm Deliveroo to scupper its workers' grassroots organising efforts, and the ongoing court case against five major tech firms for their alleged role in the maiming and deaths of people extracting raw materials in the Democratic Republic of Congo. Artificial intelligence (AI) also featured heavily in Computer Weekly's technology and ethics coverage in 2022, with stories published on the tech sector's lacklustre commitment to "ethical" AI, as well as on the pitfalls and challenges of auditing AI-powered algorithms. Police technology was another major focus of 2022, as policing bodies continue to push ahead with new tech deployments such as live facial recognition (LFR) despite serious concerns about its effectiveness, proportionality and efficacy. Other stories focused on how technology is developed and deployed, and the underlying power dynamics at play.


CES 2023: AI evolution's in a 'really important moment,' says Sony's AI ethics expert

#artificialintelligence

The viral success of OpenAI's ChatGPT has triggered discussions everywhere these days from the classroom to the boardroom. This amounts to a crucial moment for AI, Alice Xiang, Head of Sony Group's (SONY) AI Ethics Office and AI Lead Research Scientist, told Yahoo Finance Live (video above). "We're really seeing an inflection point with AI ethics, where it's going from being just something that companies are doing on their own… [and now] we're seeing policymakers really dive into this space." Xiang added: "This raises a lot of really interesting questions around how we ensure that we have governance processes in place to make sure AI that's built is compliant with relevant laws." As consumers encounter AI more frequently, and with excitement and fear, regulators are taking notice.


Digital Twin: Where do humans fit in?

arXiv.org Artificial Intelligence

Digital Twin (DT) technology is far from being comprehensive and mature, resulting in their piecemeal implementation in practice where some functions are automated by DTs, and others are still performed by humans. This piecemeal implementation of DTs often leaves practitioners wondering what roles (or functions) to allocate to DTs in a work system, and how might it impact humans. A lack of knowledge about the roles that humans and DTs play in a work system can result in significant costs, misallocation of resources, unrealistic expectations from DTs, and strategic misalignments. To alleviate this challenge, this paper answers the research question: When humans work with DTs, what types of roles can a DT play, and to what extent can those roles be automated? Specifically, we propose a two-dimensional conceptual framework, Levels of Digital Twin (LoDT). The framework is an integration of the types of roles a DT can play, broadly categorized under (1) Observer, (2) Analyst, (3) Decision Maker, and (4) Action Executor, and the extent of automation for each of these roles, divided into five different levels ranging from completely manual to fully automated. A particular DT can play any number of roles at varying levels. The framework can help practitioners systematically plan DT deployments, clearly communicate goals and deliverables, and lay out a strategic vision. A case study illustrates the usefulness of the framework.