South America
Latam-GPT: The Free, Open Source, and Collaborative AI of Latin America
Latam-GPT is new large language model being developed in and for Latin America. The project, led by the nonprofit Chilean National Center for Artificial Intelligence (CENIA), aims to help the region achieve technological independence by developing an open source AI model trained on Latin American languages and contexts. "This work cannot be undertaken by just one group or one country in Latin America: It is a challenge that requires everyone's participation," says Álvaro Soto, director of CENIA, in an interview with WIRED en Español. "Latam-GPT is a project that seeks to create an open, free, and, above all, collaborative AI model. We've been working for two years with a very bottom-up process, bringing together citizens from different countries who want to collaborate. Recently, it has also seen some more top-down initiatives, with governments taking an interest and beginning to participate in the project."
The Rosario Dataset v2: Multimodal Dataset for Agricultural Robotics
Soncini, Nicolas, Cremona, Javier, Vidal, Erica, García, Maximiliano, Castro, Gastón, Pire, Taihú
World population will grow by a third by 2050, directly impacting global food demand (Fukase and Martin (2020)). In this context, the agricultural industry should increase its production to satisfy such demand. The use of autonomous robots to carry out agricultural tasks such as seeding, harvesting, weed remotion, pest control among others is an attractive solution since it can improve the production time in a sustainable manner reducing the environmental impact and pollution. However, the implementation of autonomous robots in the agricultural field is a challenging work due the rough terrain, natural light variations, perceptual aliasing, areas with GNSS-denied signal, and the long-term robot operation required to carry out the desired applications. All these challenges cause robot localization methods to fail or perform poorly, making them impractical for real agricultural tasks, as evidenced in Cremona et al. (2022, 2023); Soncini et al. (2024); Cox et al. (2023); Bai et al. (2023); Ait et al. (2023). In the last decade, there has been a growing trend towards the creation and public availability of agricultural datasets, enabling researchers to test new techniques and develop more sophisticated algorithms to address these challenges, such as Pire et al. (2019); Kragh et al. (2017); Tanco et al. (2024). However, none of them are properly curated for evaluating multi-modal SLAM algorithms. Effective multi-modal SLAM evaluation imposes specific requirements such as hardware-synchronized sensors, 6-DOF ground-truth, and trajectories with loops to effectively test loop closure algorithms. In this work, we present a multi-modal dataset recorded by the weed removing robot developed at CIFASIS (CONICET - UNR) in a soybean agricultural field.
HSFN: Hierarchical Selection for Fake News Detection building Heterogeneous Ensemble
Coutinho, Sara B., Cruz, Rafael M. O., Nascimento, Francimaria R. S., Cavalcanti, George D. C.
Psychological biases, such as confirmation bias, make individuals particularly vulnerable to believing and spreading fake news on social media, leading to significant consequences in domains such as public health and politics. Machine learning-based fact-checking systems have been widely studied to mitigate this problem. Among them, ensemble methods are particularly effective in combining multiple classifiers to improve robustness. However, their performance heavily depends on the diversity of the constituent classifiers-selecting genuinely diverse models remains a key challenge, especially when models tend to learn redundant patterns. In this work, we propose a novel automatic classifier selection approach that prioritizes diversity, also extended by performance. The method first computes pairwise diversity between classifiers and applies hierarchical clustering to organize them into groups at different levels of granularity. A HierarchySelect then explores these hierarchical levels to select one pool of classifiers per level, each representing a distinct intra-pool diversity. The most diverse pool is identified and selected for ensemble construction from these. The selection process incorporates an evaluation metric reflecting each classifiers's performance to ensure the ensemble also generalises well. We conduct experiments with 40 heterogeneous classifiers across six datasets from different application domains and with varying numbers of classes. Our method is compared against the Elbow heuristic and state-of-the-art baselines. Results show that our approach achieves the highest accuracy on two of six datasets. The implementation details are available on the project's repository: https://github.com/SaraBCoutinho/HSFN .
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
Santos, João Guilherme Alves, Bonás, Giovana Kerche, Almeida, Thales Sales
With the growing capabilities of Large Language Models (LLMs), there is an increasing need for robust evaluation methods, especially in multilingual and non-English contexts. W e present an updated version of the BLUEX dataset, now including 2024-2025 exams and automatically generated image captions using state-of-the-art models, enhancing its relevance for data contamination studies in LLM pretraining. Captioning strategies increase accessibility to text-only models by more than 40%, producing 1,422 usable questions, more than doubling the number in the original BLUEX. W e evaluated commercial and open-source LLMs and their ability to leverage visual context through captions.
Mini Autonomous Car Driving based on 3D Convolutional Neural Networks
Moraes, Pablo, Rodriguez, Monica, Kappel, Kristofer S., Sodre, Hiago, Fernandez, Santiago, Nunes, Igor, Guterres, Bruna, Grando, Ricardo
Autonomous driving applications have become increasingly relevant in the automotive industry due to their potential to enhance vehicle safety, efficiency, and user experience, thereby meeting the growing demand for sophisticated driving assistance features. However, the development of reliable and trustworthy autonomous systems poses challenges such as high complexity, prolonged training periods, and intrinsic levels of uncertainty. Mini Autonomous Cars (MACs) are used as a practical testbed, enabling validation of autonomous control methodologies on small-scale setups. This simplified and cost-effective environment facilitates rapid evaluation and comparison of machine learning models, which is particularly useful for algorithms requiring online training. To address these challenges, this work presents a methodology based on RGB-D information and three-dimensional convolutional neural networks (3D CNNs) for MAC autonomous driving in simulated environments. We evaluate the proposed approach against recurrent neural networks (RNNs), with architectures trained and tested on two simulated tracks with distinct environmental features. Performance was assessed using task completion success, lap-time metrics, and driving consistency. Results highlight how architectural modifications and track complexity influence the models' generalization capability and vehicle control performance. The proposed 3D CNN demonstrated promising results when compared with RNNs.
AIhub monthly digest: August 2025 – causality and generative modelling, responsible multimodal AI, and IJCAI in Montréal and Guangzhou
Welcome to our monthly digest, where you can catch up with any AIhub stories you may have missed, peruse the latest news, recap recent events, and more. This month, we dive into the world of agents, learn about responsible multimodal AI, apply generative AI to computer networks, and dig into the RoboCup@Work League. This month, Sanmay Das, Tom Dietterich, Sabine Hauert, Sarit Kraus, and Michael Littman tackled the topic of agentic AI, discussing recent developments, and lessons learned from the decades of research in the autonomous agents and multiagent systems community. The 34th International Joint Conference on Artificial Intelligence (IJCAI2025) took place in Montréal from 16-22 August, with a satellite event currently being held (from 29-31 August) in Guangzhou, China. You can find out more about the programmes of both venues here, and get a flavour of what attendees got up to in our social media round-ups: Part one Part two.
Inferring processes within dynamic forest models using hybrid modeling
Pichler, Maximilian, Käber, Yannek
Modeling forest dynamics under novel climatic conditions requires a careful balance between process-based understanding and empirical flexibility. Dynamic Vegetation Models (DVM) represent ecological processes mechanistically, but their performance is prone to misspecified assumptions about functional forms. Inferring the structure of these processes and their functional forms correctly from data remains a major challenge because current approaches, such as plug-in estimators, have proven ineffective. We introduce Forest Informed Neural Networks (FINN), a hybrid modeling approach that combines a forest gap model with deep neural networks (DNN). FINN replaces processes with DNNs, which are then calibrated alongside the other mechanistic components in one unified step. In a case study on the Barro Colorado Island 50-ha plot we demonstrate that replacing the growth process with a DNN improves predictive performance and succession trajectories compared to a mechanistic version of FINN. Furthermore, we discovered that the DNN learned an ecologically plausible, improved functional form of the growth process, which we extracted from the DNN using explainable AI. In conclusion, our new hybrid modeling approach offers a versatile opportunity to infer forest dynamics from data and to improve forecasts of ecosystem trajectories under unprecedented environmental change.
Towards Trustworthy Amortized Bayesian Model Comparison
Kucharský, Šimon, Mishra, Aayush, Habermann, Daniel, Radev, Stefan T., Bürkner, Paul-Christian
Amortized Bayesian model comparison (BMC) enables fast probabilistic ranking of models via simulation-based training of neural surrogates. However, the reliability of neural surrogates deteriorates when simulation models are misspecified - the very case where model comparison is most needed. Thus, we supplement simulation-based training with a self-consistency (SC) loss on unlabeled real data to improve BMC estimates under empirical distribution shifts. Using a numerical experiment and two case studies with real data, we compare amortized evidence estimates with and without SC against analytic or bridge sampling benchmarks. SC improves calibration under model misspecification when having access to analytic likelihoods. However, it offers limited gains with neural surrogate likelihoods, making it most practical for trustworthy BMC when likelihoods are exact.
Discovering equations from data: symbolic regression in dynamical systems
Brum, Beatriz R., Lober, Luiza, Previdelli, Isolde, Rodrigues, Francisco A.
The discovery of equations from observational data is one of the fundamental pillars of the traditional scientific method. From the work of Johannes Kepler, who inferred the laws of planetary motion from meticulous astronomical observations [1] collected by Tycho Brahe [2], to Isaac Newton's theoretical formulations that consolidated classical mechanics, the process of identifying mathematical relationships underlying natural phenomena has historically been characterized by its manual nature, based essentially on systematic trial-and-error procedures. However, in recent decades, the advent of Big Data, characterized by the production of an immense volume of complex, mostly nonlinear, data, in several fields has driven a new search for physical laws. Faced with the need to analyze these data sets to understand their intrinsic structure and derive symbolic representations that capture the integral behavior of a system, the demand for advanced analytical methods has become growing and indispensable. With the emergence of modern computational techniques, this process has undergone a radical transformation, driving the widespread development and use of various regression techniques.
Unbiased Stochastic Optimization for Gaussian Processes on Finite Dimensional RKHS
Current methods for stochastic hyperparameter learning in Gaussian Processes (GPs) rely on approximations, such as computing biased stochastic gradients or using inducing points in stochastic variational inference. However, when using such methods we are not guaranteed to converge to a stationary point of the true marginal likelihood. In this work, we propose algorithms for exact stochastic inference of GPs with kernels that induce a Reproducing Kernel Hilbert Space (RKHS) of moderate finite dimension. Our approach can also be extended to infinite dimensional RKHSs at the cost of forgoing exactness. Both for finite and infinite dimensional RKHSs, our method achieves better experimental results than existing methods when memory resources limit the feasible batch size and the possible number of inducing points.