safegraph
Inferring Dynamic Networks from Marginals with Iterative Proportional Fitting
Chang, Serina, Koehler, Frederic, Qu, Zhaonan, Leskovec, Jure, Ugander, Johan
A common network inference problem, arising from real-world data constraints, is how to infer a dynamic network from its time-aggregated adjacency matrix and time-varying marginals (i.e., row and column sums). Prior approaches to this problem have repurposed the classic iterative proportional fitting (IPF) procedure, also known as Sinkhorn's algorithm, with promising empirical results. However, the statistical foundation for using IPF has not been well understood: under what settings does IPF provide principled estimation of a dynamic network from its marginals, and how well does it estimate the network? In this work, we establish such a setting, by identifying a generative network model whose maximum likelihood estimates are recovered by IPF. Our model both reveals implicit assumptions on the use of IPF in such settings and enables new analyses, such as structure-dependent error bounds on IPF's parameter estimates. When IPF fails to converge on sparse network data, we introduce a principled algorithm that guarantees IPF converges under minimal changes to the network structure. Finally, we conduct experiments with synthetic and real-world data, which demonstrate the practical value of our theoretical and algorithmic contributions.
BiasBuster: a Neural Approach for Accurate Estimation of Population Statistics using Biased Location Data
Zeighami, Sepanta, Shahabi, Cyrus
While extremely useful (e.g., for COVID-19 forecasting and policy-making, urban mobility analysis and marketing, and obtaining business insights), location data collected from mobile devices often contain data from a biased population subset, with some communities over or underrepresented in the collected datasets. As a result, aggregate statistics calculated from such datasets (as is done by various companies including Safegraph, Google, and Facebook), while ignoring the bias, leads to an inaccurate representation of population statistics. Such statistics will not only be generally inaccurate, but the error will disproportionately impact different population subgroups (e.g., because they ignore the underrepresented communities). This has dire consequences, as these datasets are used for sensitive decision-making such as COVID-19 policymaking. This paper tackles the problem of providing accurate population statistics using such biased datasets. We show that statistical debiasing, although in some cases useful, often fails to improve accuracy. We then propose BiasBuster, a neural network approach that utilizes the correlations between population statistics and location characteristics to provide accurate estimates of population statistics. Extensive experiments on real-world data show that BiasBuster improves accuracy by up to 2 times in general and up to 3 times for underrepresented populations.
Toyota Leaked Vehicle Data of 2 Million Customers
SafeGraph, the data broker famous for selling location data linked to abortion clinic visits, is now a US military contractor. Documents obtained by WIRED reveal that the company landed an initial contract with the US Air Force and is hoping the Pentagon will buy a tool that SafeGraph says will pinpoint locations not to bomb, like schools and hospitals. Your data is, of course, everywhere--likely including in the training data of generative AI tools like ChatGPT. Fortunately, at least some users can request that OpenAI, which created the tool, delete their data. It's also possible to delete your chat history with ChatGPT.
Five insights about harnessing data and AI from leaders at the frontier
What was once unknowable can now be quickly discovered with a few queries. Decision makers no longer have to rely on gut instinct; today they have more extensive and precise evidence at their fingertips. New sources of data, fed into systems powered by machine learning and AI, are at the heart of this transformation. The information flowing through the physical world and the global economy is staggering in scope. It comes from thousands of sources: sensors, satellite imagery, web traffic, digital apps, videos, and credit card transactions, just to name a few. These types of data can transform decision making.
Location data analytics provider SafeGraph raises $45M
SafeGraph, a startup using AI to create and maintain mobility datasets, today announced that it raised $45 million in funding led by Sapphire Ventures. With the investment, SafeGraph plans to capitalize on the expanding international market of data buyers and offer new ways for companies to buy data through its network. Location data is fast-becoming a hot commodity. In May 2019, 94% of mobile marketers in the U.S. surveyed by Statista said they were already using location data for advertising purposes, while 94% indicated that they were planning to do so in the future. While more users worldwide are refusing to share location data with apps, marketers say the data is incredibly valuable for targeting purposes and personalizing customer experiences.
A Non-Technical Introduction to Machine Learning โ SafeGraph
Machine learning is a field that threatens to both augment and undermine exactly what it means to be human, and it's becoming increasingly important that you--yes, you--actually understand it. I don't think you should need to have a technical background to know what machine learning is or how it's done. Too much of the discussion about this field is either too technical or too uninformed, and, through this blog, I hope to level the playing field. This is for smart, ambitious people who want to know more about machine learning but who don't care about the esoteric statistical and computational details underlying the field. You don't need to know any math, statistics, or computer science to read and understand it.