Goto

Collaborating Authors

 Education


Equality of opportunity in travel behavior prediction with deep neural networks and discrete choice models

arXiv.org Machine Learning

Although researchers increasingly adopt machine learning to model travel behavior, they predominantly focus on prediction accuracy, ignoring the ethical challenges embedded in machine learning algorithms. This study introduces an important missing dimension - computational fairness - to travel behavior analysis. We first operationalize computational fairness by equality of opportunity, then differentiate between the bias inherent in data and the bias introduced by modeling. We then demonstrate the prediction disparities in travel behavior modeling using the 2017 National Household Travel Survey (NHTS) and the 2018-2019 My Daily Travel Survey in Chicago. Empirically, deep neural network (DNN) and discrete choice models (DCM) reveal consistent prediction disparities across multiple social groups: both over-predict the false negative rate of frequent driving for the ethnic minorities, the low-income and the disabled populations, and falsely predict a higher travel burden of the socially disadvantaged groups and the rural populations than reality. Comparing DNN with DCM, we find that DNN can outperform DCM in prediction disparities because of DNN's smaller misspecification error. To mitigate prediction disparities, this study introduces an absolute correlation regularization method, which is evaluated with synthetic and real-world data. The results demonstrate the prevalence of prediction disparities in travel behavior modeling, and the disparities still persist regarding a variety of model specifics such as the number of DNN layers, batch size and weight initialization. Since these prediction disparities can exacerbate social inequity if prediction results without fairness adjustment are used for transportation policy making, we advocate for careful consideration of the fairness problem in travel behavior modeling, and the use of bias mitigation algorithms for fair transport decisions.


Finetuning Transformer Models to Build ASAG System

arXiv.org Artificial Intelligence

Research towards creating systems for automatic grading of student answers to quiz and exam questions in educational settings has been ongoing since 1966. Over the years, the problem was divided into many categories. Among them, grading text answers were divided into short answer grading, and essay grading. The goal of this work was to develop an ML-based short answer grading system. I hence built a system which uses finetuning on Roberta Large Model pretrained on STS benchmark dataset and have also created an interface to show the production readiness of the system. I evaluated the performance of the system on the Mohler extended dataset and SciEntsBank Dataset. The developed system achieved a Pearsons Correlation of 0.82 and RMSE of 0.7 on the Mohler Dataset which beats the SOTA performance on this dataset which is correlation of 0.805 and RMSE of 0.793. Additionally, Pearsons Correlation of 0.79 and RMSE of 0.56 was achieved on the SciEntsBank Dataset, which only reconfirms the robustness of the system. A few observations during achieving these results included usage of batch size of 1 produced better results than using batch size of 16 or 32 and using huber loss as loss function performed well on this regression task. The system was tried and tested on train and validation splits using various random seeds and still has been tweaked to achieve a minimum of 0.76 of correlation and a maximum 0.15 (out of 1) RMSE on any dataset.


Researcher/Lecturer Revisiting photogrammetry using deep learning

#artificialintelligence

You will establish a research line at the interface between photogrammetry and deep learning developing innovative solutions. You will explore how to embed deep learning algorithms in the image orientation and 3D reconstruction tasks and their possible fusion with semantic segmentation. Your goal will be to implement reliable and robust solutions to be applied in several applications, with specific regard to the use of very high-resolution data. Your background is in deep learning, photogrammetry and computer vision and must be confirmed by an excellent track of record. Besides your research activities, you may be occasionally involved in educational activities.


Generally Intelligent #12: Jacob Steinhardt, UC Berkeley, on machine learning safety, alignment and measurement

#artificialintelligence

Jacob Steinhardt (Google Scholar) (Website) is an assistant professor at UC Berkeley. His main research interest is in designing machine learning systems that are reliable and aligned with human values. Some of his specific research directions include robustness, rewards specification and reward hacking, as well as scalable alignment. His most recent paper at ICLR 2021 proposes a new test to measure an NLP model's accuracy on a wide variety of tasks, ranging from mathematics, US history, law, and more. It provides a measurement tool to help researchers specify an important problem: while current models can achieve superhuman performance on benchmarks, they lack the ability to understand language on a whole. Another of Jacob's papers at ICLR focuses on measuring a language model's knowledge of basic concepts of morality. It shows that current language models have a promising but incomplete ability to predict basic human ethical judgements. "Test accuracy is a very limited metric." "You might not be able to get lots of feedback on human values." Below are the show notes and full transcript. As always, please feel free to reach out with feedback, ideas, and questions! I think it required me to learn to become a significantly better writer. And I think that helped later on, because it made me feel more comfortable pursuing unusual ideas. I knew I had the skills to present those ideas. As long as I believed in them, I could get other people to believe in them." You just want this very diverse distribution of things that are deeply ingrained in evolutionary history as opposed to being part of explicit reasoning" First of all, test accuracy is a very limited metric. What are we trying to do with it? For a while, there was a lot of climate skepticism or climate denial. At some point it becomes pretty clear, when there's regular heat waves fires and that sort of thing. You probably wanted to do something about it before that point. Having these more subtle measurements that you can look at are important. And the other thing is I think it actually laid the groundwork for the more extreme weather events to become a convincing signal. Jacob Steinhardt: Another thing that I'm interested in is just measuring the progress in capabilities, getting different AI capabilities seems important. Vision tasks just seem to be falling like flies. I don't know if there's any vision tasks that's survived for more than a year and a few tasks seem a little bit better, but I think those are also starting to fall like flies. I know we've come up with a few harder tasks. ML Systems are still not very good at math. Humans also aren't very good at math, but also not good at law it turns out.


Caldera chronicles: Computers taught to recognize Yellowstone quakes

#artificialintelligence

Yellowstone Caldera Chronicles is a weekly column written by scientists and collaborators of the Yellowstone Volcano Observatory. This week's contribution is from Keith Koper, director of the University of Utah Seismograph Stations and professor at the University of Utah Department of Geology and Geophysics, and Alysha Armstrong, graduate student at the University of Utah Department of Geology and Geophysics. While the automated monitoring system currently in place for detecting and processing earthquakes in Yellowstone National Park works well most of the time, its solutions need to be reviewed and refined by a seismic analyst. This means that the larger earthquakes -- generally over M1-- get most of the attention, and smaller earthquakes, which are harder to locate, are not always processed. The current system can also struggle in situations like earthquake swarms, where there is a lot of seismicity close together in space and time.


The things we (ought to) learn…

#artificialintelligence

Yester night my drum tutor taught me a new drum-roll. Just having to grasp that new technique sparked within me an introspective tangent, as usually is the case. For close to one and a half years' now, he has been patiently taking me through the plenty foundations and a few advanced techniques in drumming. All thanks to him, I have seen how I have progressed from a wannabe drummer to at least a basic one. For me, that continual process has been in many regards, nothing short of revelatory.


5 Best Online Biostatistics Programs and Courses

#artificialintelligence

Are you looking for Best Online Biostatistics Programs and Courses?… If yes, then your search will end here. In this article, I am going to share the 5 Best Online Biostatistics Programs and Courses with you. So, give your few minutes to this article and find out the best online Biostatistics program for you. The goal of Biostatistics is to advance statistical science and its application to problems of human health and disease, with the ultimate goal of advancing the public's health.


Cal Poly Project Leverages Artificial Intelligence Deep Learning to Aid Wildfire Recovery

#artificialintelligence

SAN LUIS OBISPO –– A pair of Cal Poly professors and a team of students have used artificial intelligence to train a computer to quickly assess wildfire damage -- potentially improving response time for efforts to recover from major wildfires. Accurate and timely damage assessment has become critical for response and recovery as the threat of wildfires increases. Damage assessment reports inform first responders' strategies, affect residents' ability to file insurance claims, and guide state and federal authorities' plans for future disaster relief and financial aid. To date, most wildfire event inspectors must personally visit affected areas and manually document the severity of building damage, a process that often takes weeks. Social sciences Assistant Professor Andrew Fricker, computer science Assistant Professor Jonathan Ventura, visiting Cal Poly undergraduate student Gustave Rousselet, and a team of Stanford doctoral students sought to streamline this process with artificial intelligence (AI) deep learning.


GitHub - ahmedbahaaeldin/From-0-to-Research-Scientist-resources-guide: Detailed and tailored guide for undergraduate students or anybody want to dig deep into the field of AI with solid foundation.

#artificialintelligence

This guide is designated to anybody with basic programming knowledge or a computer science background interested in becoming a Research Scientist with on Deep Learning and NLP. You can go Bottom-Up or Top-Down both works well and it is actually crucial to know which approach suites you the best. If you are okay with studying lots of mathematical concepts without application then use Bottom-Up. If you want to go hands-on first then use the Top-Down first. The Mathematical Foundation part is for all Artificial Intelligence branches such as Machine Learning, Reinforcement Learning, Computer Vision and so on. AI is heavily math-theory based so a solid foundation is essential.


Sinkhorn Distributionally Robust Optimization

arXiv.org Machine Learning

Decision-making problems under uncertainty have broad applications in operations research, machine learning, engineering, and economics. When the data involves uncertainty due to measurement error, insufficient sample size, contamination, and anomalies, or model misspecification, distributionally robust optimization (DRO) is a promising approach to data-driven optimization, by seeking a minimax robust optimal decision that minimizes the expected loss under the most adverse distribution within a given set of relevant distributions, called ambiguity set. It provides a principled framework to produce a solution with more promising out-of-sample performance than the traditional sample average approximation (SAA) method for stochastic programming [86]. We refer to [81] for a recent survey on DRO. At the core of DRO is the choice of the ambiguity set. Ideally, a good ambiguity set should take account of the properties of practical applications while maintaining the computational tractability of resulted DRO formulation; and it should be rich enough to contain all distributions relevant to the decision-making but, at the same time, should not include unnecessary distributions that lead to overly conservative decisions. Various DRO formulations have been proposed in the literature. Among them, the ambiguity set based on Wasserstein distance has recently received much attention [104, 67, 17, 46]. The Wasserstein distance incorporates the geometry of sample space, and thereby is suitable for comparing distributions with non-overlapping supports and hedging against data perturbations [46].