street-level imagery
Unsupervised Urban Tree Biodiversity Mapping from Street-Level Imagery Using Spatially-Aware Visual Clustering
Abuhani, Diaa Addeen, Seccaroni, Marco, Mazzarello, Martina, Zualkernan, Imran, Duarte, Fabio, Ratti, Carlo
Urban tree biodiversity is critical for climate resilience, ecological stability, and livability in cities, yet most municipalities lack detailed knowledge of their canopies. Field-based inventories provide reliable estimates of Shannon and Simpson diversity but are costly and time-consuming, while supervised AI methods require labeled data that often fail to generalize across regions. We introduce an unsupervised clustering framework that integrates visual embeddings from street-level imagery with spatial planting patterns to estimate biodiversity without labels. Applied to eight North American cities, the method recovers genus-level diversity patterns with high fidelity, achieving low Wasserstein distances to ground truth for Shannon and Simpson indices and preserving spatial autocorrelation. This scalable, fine-grained approach enables biodiversity mapping in cities lacking detailed inventories and offers a pathway for continuous, low-cost monitoring to support equitable access to greenery and adaptive management of urban ecosystems.
Open-source maps should help driverless cars navigate our cities more safely
Our current street maps aren't much good for helping driverless cars get around. Although we've mapped most roads, they get updated only every couple of years. And these maps don't log any roadside infrastructure such as road signs, driveways, and lane markings. Without this extra layer of information, it will be much harder to get autonomous cars to navigate our cities safely. Robotic deliveries, too, will eventually require precise details of road surfaces, sidewalks, and obstacles.
How we find and update Points of Interest from street-level imagery
The second step is "text recognition". "Using deep learning," Xing continues, "we classify text and filter out any "noise", such as windows or patterns, for example, which are sometimes misclassified as text. Our recurrent neural network uses a sequence model to transcribe the text, with an attention model focusing on specific parts of the images to better recognize individual characters, numbers and symbols."