South America
A Comparative Study on Crime in Denver City Based on Machine Learning and Data Mining
To ensure the security of the general mass, crime prevention is one of the most higher priorities for any government. An accurate crime prediction model can help the government, law enforcement to prevent violence, detect the criminals in advance, allocate the government resources, and recognize problems causing crimes. To construct any future-oriented tools, examine and understand the crime patterns in the earliest possible time is essential. In this paper, I analyzed a real-world crime and accident dataset of Denver county, USA, from January 2014 to May 2019, which containing 478,578 incidents. This project aims to predict and highlights the trends of occurrence that will, in return, support the law enforcement agencies and government to discover the preventive measures from the prediction rates. At first, I apply several statistical analysis supported by several data visualization approaches. Then, I implement various classification algorithms such as Random Forest, Decision Tree, AdaBoost Classifier, Extra Tree Classifier, Linear Discriminant Analysis, K-Neighbors Classifiers, and 4 Ensemble Models to classify 15 different classes of crimes. The outcomes are captured using two popular test methods: train-test split, and k-fold cross-validation. Moreover, to evaluate the performance flawlessly, I also utilize precision, recall, F1-score, Mean Squared Error (MSE), ROC curve, and paired-T-test. Except for the AdaBoost classifier, most of the algorithms exhibit satisfactory accuracy. Random Forest, Decision Tree, Ensemble Model 1, 3, and 4 even produce me more than 90% accuracy. Among all the approaches, Ensemble Model 4 presented superior results for every evaluation basis. This study could be useful to raise the awareness of peoples regarding the occurrence locations and to assist security agencies to predict future outbreaks of violence in a specific area within a particular time.
Learning Generative Models using Denoising Density Estimators
Bigdeli, Siavash A., Lin, Geng, Portenier, Tiziano, Dunbar, L. Andrea, Zwicker, Matthias
Learning generative probabilistic models that can estimate the continuous density given a set of samples, and that can sample from that density, is one of the fundamental challenges in unsupervised machine learning. In this paper we introduce a new approach to obtain such models based on what we call denoising density estimators (DDEs). A DDE is a scalar function, parameterized by a neural network, that is efficiently trained to represent a kernel density estimator of the data. Leveraging DDEs, our main contribution is to develop a novel approach to obtain generative models that sample from given densities. We prove that our algorithms to obtain both DDEs and generative models are guaranteed to converge to the correct solutions. Advantages of our approach include that we do not require specific network architectures like in normalizing flows, ordinary differential equation solvers as in continuous normalizing flows, nor do we require adversarial training as in generative adversarial networks (GANs). Finally, we provide experimental results that demonstrate practical applications of our technique.
A Correspondence Analysis Framework for Author-Conference Recommendations
Iyer, Rahul Radhakrishnan, Sharma, Manish, Saradhi, Vijaya
For many years, achievements and discoveries made by scientists are made aware through research papers published in appropriate journals or conferences. Often, established scientists and especially newbies are caught up in the dilemma of choosing an appropriate conference to get their work through. Every scientific conference and journal is inclined towards a particular field of research and there is a vast multitude of them for any particular field. Choosing an appropriate venue is vital as it helps in reaching out to the right audience and also to further one's chance of getting their paper published. In this work, we address the problem of recommending appropriate conferences to the authors to increase their chances of acceptance. We present three different approaches for the same involving the use of social network of the authors and the content of the paper in the settings of dimensionality reduction and topic modeling. In all these approaches, we apply Correspondence Analysis (CA) to derive appropriate relationships between the entities in question, such as conferences and papers. Our models show promising results when compared with existing methods such as content-based filtering, collaborative filtering and hybrid filtering.
Technology Trends of 2020
At the Last Futurist, we enjoy looking at AI Trends and digital transformation trends. In between those two are more broad technology trends. In fact these topics make up the mission statement of this new news site. However the last decade had a lot of technology and gadgets that didn't fare so well in the real world. The decade was mobile all the way, with mass adoption taking place the way we might expect the brain-computer interface (BCI) to achieve mass adoption in a future decade years from now. In the decade ahead the move to automated stores and electric vehicles are real trends, but it's important to differentiate the hype from the reality. Autonomous vehicles, quantum computing going mainstream, better self-learning AI, hang on a second! Even mass adoption of digital currencies is coming faster. From computers to the internet and smart phones, a few generations shows a lot of progress. But technology never stands still. Advertising has scaled a world of surveillance capitalism normalization and an AI-arms race is now taking place. Most technology trends and AI listicles only touch the surface of how humans are embedding technology increasingly into their lives. However looking at it from the perspectives of many industries and across technology and innovation stacks gives a more complete picture. The real world and customer experience are the real tests for new technological innovations and pivots. It will take decades for 3D printing, quantum computing and an AGI to even become mature, but an age of biotechnology and AI in healthcare, education and finance is inevitable. From Huawei, to ByteDance (TikTok), to Didi, China will wage major battles for global market share in 5G, consumer apps, E-commerce, mobile payments and ride sharing, among others. Chinese led tech companies -- with the support of the Chinese Government and venture funds such as Softbank Vision Fund -- can mean that in the 2020s China's ecosystem fully replaces Silicon Valley as the leader of innovation. In 2019, some believe this has already occurred.
Argentina boosts security at airports, U.S. Embassy over Iran tensions
BUENOS AIRES – Argentina's government boosted security at its airports, borders and the U.S. Embassy in Buenos Aires as tensions simmer between the United States and Iran, the South American country's defense minister told local media on Monday. Argentina, which suffered two attacks, in 1992 and 1994, decided to raise its alert level days after a U.S. drone strike killed Iranian military commander Qassem Soleimani in Iraq, stoking global fears of retaliation attacks. "Because of the history of two attacks we had, Argentina must be on alert for this type of conflict worldwide," Defense Minister Agustin Rossi told local news site Infobae. More than 100 people were killed in two attacks in Argentina in the 1990s. In 1992, the Israeli Embassy in Buenos Aires was attacked with a car bomb, killing 29 people.
Poly-time universality and limitations of deep learning
The goal of this paper is to characterize function distributions that deep learning can or cannot learn in poly-time. A universality result is proved for SGD-based deep learning and a non-universality result is proved for GD-based deep learning; this also gives a separation between SGD-based deep learning and statistical query algorithms: (1) {\it Deep learning with SGD is efficiently universal.} Any function distribution that can be learned from samples in poly-time can also be learned by a poly-size neural net trained with SGD on a poly-time initialization with poly-steps, poly-rate and possibly poly-noise. Therefore deep learning provides a universal learning paradigm: it was known that the approximation and estimation errors could be controlled with poly-size neural nets, using ERM that is NP-hard; this new result shows that the optimization error can also be controlled with SGD in poly-time. The picture changes for GD with large enough batches: (2) {\it Result (1) does not hold for GD:} Neural nets of poly-size trained with GD (full gradients or large enough batches) on any initialization with poly-steps, poly-range and at least poly-noise cannot learn any function distribution that has super-polynomial {\it cross-predictability,} where the cross-predictability gives a measure of ``average'' function correlation -- relations and distinctions to the statistical dimension are discussed. In particular, GD with these constraints can learn efficiently monomials of degree $k$ if and only if $k$ is constant. Thus (1) and (2) point to an interesting contrast: SGD is universal even with some poly-noise while full GD or SQ algorithms are not (e.g., parities).
Stochastic Weight Averaging in Parallel: Large-Batch Training that Generalizes Well
Gupta, Vipul, Serrano, Santiago Akle, DeCoste, Dennis
We propose Stochastic Weight Averaging in Parallel (SW AP), an algorithm to accelerate DNN training. Our algorithm uses large mini-batches to compute an approximate solution quickly and then refines it by averaging the weights of multiple models computed independently and in parallel. The resulting models generalize equally well as those trained with small mini-batches but are produced in a substantially shorter time. We demonstrate the reduction in training time and the good generalization performance of the resulting models on the computer vision datasets CIFAR10, CIFAR100, and ImageNet. Stochastic gradient descent (SGD) and its variants are the de-facto methods to train deep neural networks (DNNs). Each iteration of SGD computes an estimate of the objective's gradient by sampling a mini-batch of the available training data and computing the gradient of the loss restricted to the sampled data. A popular strategy to accelerate DNN training is to increase the mini-batch size together with the available computational resources. Larger mini-batches produce more precise gradient estimates; these allow for higher learning rates and achieve larger reductions of the training loss per iteration.
Context-Aware Design of Cyber-Physical Human Systems (CPHS)
Mukhopadhyay, Supratik, Liu, Qun, Collier, Edward, Zhu, Yimin, Gudishala, Ravindra, Chokwitthaya, Chanachok, DiBiano, Robert, Nabijiang, Alimire, Saeidi, Sanaz, Sidhanta, Subhajit, Ganguly, Arnab
Recently, it has been widely accepted by the research community that interactions between humans and cyber-physical infrastructures have played a significant role in determining the performance of the latter. The existing paradigm for designing cyber-physical systems for optimal performance focuses on developing models based on historical data. The impacts of context factors driving human system interaction are challenging and are difficult to capture and replicate in existing design models. As a result, many existing models do not or only partially address those context factors of a new design owing to the lack of capabilities to capture the context factors. This limitation in many existing models often causes performance gaps between predicted and measured results. We envision a new design environment, a cyber-physical human system (CPHS) where decision-making processes for physical infrastructures under design are intelligently connected to distributed resources over cyberinfrastructure such as experiments on design features and empirical evidence from operations of existing instances. The framework combines existing design models with context-aware design-specific data involving human-infrastructure interactions in new designs, using a machine learning approach to create augmented design models with improved predictive powers.
Global LegalTech Artificial Intelligence Market: Dynamic Business Environment – Food & Beverage Herald
The "LegalTech Artificial Intelligence Market" is evolving at an exciting pace driven by changing dynamics and risk ecosystem, an analysis of which forms the crux of the report. The study on the global LegalTech Artificial Intelligence Market takes a closer look at several regional trends and the emerging regulatory landscape to assess its prospects. The critical evaluation of the various growth factors and opportunities in the global LegalTech Artificial Intelligence Market offered in the analyses helps in assessing the lucrativeness of its key segments. Summary of Market: The global LegalTech Artificial Intelligence market is valued at xx million US$ in 2019 is expected to reach xx million US$ by the end of 2025, growing at a CAGR of xx% during 2019-2025. Legal technology, also known asLegal Tech, refers to the use oftechnologyandsoftwareto providelegal services.
Topic Extraction of Crawled Documents Collection using Correlated Topic Model in MapReduce Framework
The tremendous increase in the amount of available research documents impels researchers to propose topic models to extract the latent semantic themes of a documents collection. However, how to extract the hidden topics of the documents collection has become a crucial task for many topic model applications. Moreover, conventional topic modeling approaches suffer from the scalability problem when the size of documents collection increases. In this paper, the Correlated Topic Model with variational Expectation-Maximization algorithm is implemented in MapReduce framework to solve the scalability problem. The proposed approach utilizes the dataset crawled from the public digital library. In addition, the full-texts of the crawled documents are analysed to enhance the accuracy of MapReduce CTM. The experiments are conducted to demonstrate the performance of the proposed algorithm. From the evaluation, the proposed approach has a comparable performance in terms of topic coherences with LDA implemented in MapReduce framework.