Goto

Collaborating Authors

 South America


On the Identifiability of Tensor Ranks via Prior Predictive Matching

arXiv.org Machine Learning

Selecting the latent dimensions (ranks) in tensor factorization is a central challenge that often relies on heuristic methods. This paper introduces a rigorous approach to determine rank identifiability in probabilistic tensor models, based on prior predictive moment matching. We transform a set of moment matching conditions into a log-linear system of equations in terms of marginal moments, prior hyperparameters, and ranks; establishing an equivalence between rank identifiability and the solvability of such system. We apply this framework to four foundational tensor-models, demonstrating that the linear structure of the PARAFAC/CP model, the chain structure of the Tensor Train model, and the closed-loop structure of the Tensor Ring model yield solvable systems, making their ranks identifiable. In contrast, we prove that the symmetric topology of the Tucker model leads to an underdetermined system, rendering the ranks unidentifiable by this method. For the identifiable models, we derive explicit closed-form rank estimators based on the moments of observed data only. We empirically validate these estimators and evaluate the robustness of the proposal.


deFOREST: Fusing Optical and Radar satellite data for Enhanced Sensing of Tree-loss

arXiv.org Machine Learning

In this paper we develop a deforestation detection pipeline that incorporates optical and Synthetic Aperture Radar (SAR) data. A crucial component of the pipeline is the construction of anomaly maps of the optical data, which is done using the residual space of a discrete Karhunen-Loève (KL) expansion. Anomalies are quantified using a concentration bound on the distribution of the residual components for the nominal state of the forest. This bound does not require prior knowledge on the distribution of the data. This is in contrast to statistical parametric methods that assume knowledge of the data distribution, an impractical assumption that is especially infeasible for high dimensional data such as ours. Once the optical anomaly maps are computed they are combined with SAR data, and the state of the forest is classified by using a Hidden Markov Model (HMM). We test our approach with Sentinel-1 (SAR) and Sentinel-2 (Optical) data on a $92.19\,km \times 91.80\,km$ region in the Amazon forest. The results show that both the hybrid optical-radar and optical only methods achieve high accuracy that is superior to the recent state-of-the-art hybrid method. Moreover, the hybrid method is significantly more robust in the case of sparse optical data that are common in highly cloudy regions.


Interaction Concordance Index: Performance Evaluation for Interaction Prediction Methods

arXiv.org Machine Learning

Consider two sets of entities and their members' mutual affinity values, say drug-target affinities (DTA). Drugs and targets are said to interact in their effects on DTAs if drug's effect on it depends on the target. Presence of interaction implies that assigning a drug to a target and another drug to another target does not provide the same aggregate DTA as the reversed assignment would provide. Accordingly, correctly capturing interactions enables better decision-making, for example, in allocation of limited numbers of drug doses to their best matching targets. Learning to predict DTAs is popularly done from either solely from known DTAs or together with side information on the entities, such as chemical structures of drugs and targets. In this paper, we introduce interaction directions' prediction performance estimator we call interaction concordance index (IC-index), for both fixed predictors and machine learning algorithms aimed for inferring them. IC-index complements the popularly used DTA prediction performance estimators by evaluating the ratio of correctly predicted directions of interaction effects in data. First, we show the invariance of IC-index on predictors unable to capture interactions. Secondly, we show that learning algorithm's permutation equivariance regarding drug and target identities implies its inability to capture interactions when either drug, target or both are unseen during training. In practical applications, this equivariance is remedied via incorporation of appropriate side information on drugs and targets. We make a comprehensive empirical evaluation over several biomedical interaction data sets with various state-of-the-art machine learning algorithms. The experiments demonstrate how different types of affinity strength prediction methods perform in terms of IC-index complementing existing prediction performance estimators.


From Guess2Graph: When and How Can Unreliable Experts Safely Boost Causal Discovery in Finite Samples?

arXiv.org Artificial Intelligence

Causal discovery algorithms often perform poorly with limited samples. While integrating expert knowledge (including from LLMs) as constraints promises to improve performance, guarantees for existing methods require perfect predictions or uncertainty estimates, making them unreliable for practical use. We propose the Guess2Graph (G2G) framework, which uses expert guesses to guide the sequence of statistical tests rather than replacing them. This maintains statistical consistency while enabling performance improvements. We develop two instantiations of G2G: PC-Guess, which augments the PC algorithm, and gPC-Guess, a learning-augmented variant designed to better leverage high-quality expert input. Theoretically, both preserve correctness regardless of expert error, with gPC-Guess provably outperforming its non-augmented counterpart in finite samples when experts are "better than random."


Do Large Language Models Show Biases in Causal Learning? Insights from Contingency Judgment

arXiv.org Artificial Intelligence

Causal learning is the cognitive process of developing the capability of making causal inferences based on available information, often guided by normative principles. This process is prone to errors and biases, such as the illusion of causality, in which people perceive a causal relationship between two variables despite lacking supporting evidence. This cognitive bias has been proposed to underlie many societal problems, including social prejudice, stereotype formation, misinformation, and superstitious thinking. In this work, we examine whether large language models are prone to developing causal illusions when faced with a classic cognitive science paradigm: the contingency judgment task. To investigate this, we constructed a dataset of 1,000 null contingency scenarios (in which the available information is not sufficient to establish a causal relationship between variables) within medical contexts and prompted LLMs to evaluate the effectiveness of potential causes. Our findings show that all evaluated models systematically inferred unwarranted causal relationships, revealing a strong susceptibility to the illusion of causality. While there is ongoing debate about whether LLMs genuinely understand causality or merely reproduce causal language without true comprehension, our findings support the latter hypothesis and raise concerns about the use of language models in domains where accurate causal reasoning is essential for informed decision-making.


Quechua Speech Datasets in Common Voice: The Case of Puno Quechua

arXiv.org Artificial Intelligence

Under-resourced languages, such as Quechuas, face data and resource scarcity, hindering their development in speech technology. To address this issue, Common Voice presents a crucial opportunity to foster an open and community-driven speech dataset creation. This paper examines the integration of Quechua languages into Common Voice. We detail the current 17 Quechua languages, presenting Puno Quechua (ISO 639-3: qxp) as a focused case study that includes language onboarding and corpus collection of both reading and spontaneous speech data. Our results demonstrate that Common Voice now hosts 191.1 hours of Quechua speech (86\% validated), with Puno Quechua contributing 12 hours (77\% validated), highlighting the Common Voice's potential. We further propose a research agenda addressing technical challenges, alongside ethical considerations for community engagement and indigenous data sovereignty. Our work contributes towards inclusive voice technology and digital empowerment of under-resourced language communities.


Sam Fender wins 2025 Mercury Prize for album of the year

BBC News

Sam Fender has won the 2025 Mercury Prize for his third album, People Watching, a steely-eyed dissection of working-class life in the north of England. The singer looked stunned when his name was announced. I didn't think that was going to happen at all, he told the BBC as he came off stage. I've spent the last 10 minutes crying. Fender beat the likes of Pulp and Wolf Alice - both former winners of the £25,000 prize for the best British or Irish album of the year - at a star-studded ceremony in Newcastle's Utilita Arena.


Trump says he will meet Putin in Hungary for Ukraine talks after 'very productive' call

BBC News

Trump says he will meet Putin in Hungary for Ukraine talks after'very productive' call US President Donald Trump says great progress was made during a phone call with Russian President Vladimir Putin on Thursday, with the pair agreeing to face-to-face talks in Hungary. He said the call, the first with Putin since mid-August, was very productive, adding that teams from Washington and Moscow will meet next week. Trump did not confirm a date for his meeting with Putin in Budapest. The Kremlin said work on the summit would begin immediately after the extremely frank and trustful call. The talks came a day before Ukraine's President Zelensky was to visit the White House, and with Trump weighing whether to arm Ukraine with Tomahawk missiles capable of striking deep into Russia.


Adiós, AirPods

The Atlantic - Technology

Apple promises to put an AI interpreter in everyone's ears. It couldn't even help me order tamales. Earlier this week, I stopped for breakfast in Sunset Park, Brooklyn, a largely Hispanic neighborhood where street vendors sell tamales and rice pudding out of orange Gatorade coolers. I speak some Spanish, but I wanted to test out Apple's new "Live Translation" feature, which has been advertised as a sort of interpreter in your ears. I popped in my AirPods, pulled up the Translate app, and approached.


EU sets 2027 target for anti-drone system to defend against Russia

BBC News

EU foreign policy chief Kaja Kallas has said a new anti-drone system should be fully operational by the end of 2027, as part of a drive to toughen defences against Russia and be fully prepared for possible conflict by 2030. Drones are already redefining warfare. Having drone defences is no longer optional for anyone, Kallas said, referring to Russia's ongoing war in Ukraine and fears that Moscow may attack the EU. The European Commission's defence roadmap also proposes strengthening the EU's eastern borders and building air and space shields. Several EU nations have faced Russian incursions into their airspace and US President Donald Trump has urged the bloc to do more to defend itself.