Industry
Interpretable Model-Aware Counterfactual Explanations for Random Forest
Harvey, Joshua S., Feng, Guanchao, Meesala, Sai Anusha, Zhao, Tina, Mehta, Dhagash
Despite their enormous predictive power, machine learning models are often unsuitable for applications in regulated industries such as finance, due to their limited capacity to provide explanations. While model-agnostic frameworks such as Shapley values have proved to be convenient and popular, they rarely align with the kinds of causal explanations that are typically sought after. Counterfactual case-based explanations, where an individual is informed of which circumstances would need to be different to cause a change in outcome, may be more intuitive and actionable. However, finding appropriate counterfactual cases is an open challenge, as is interpreting which features are most critical for the change in outcome. Here, we pose the question of counterfactual search and interpretation in terms of similarity learning, exploiting the representation learned by the random forest predictive model itself. Once a counterfactual is found, the feature importance of the explanation is computed as a function of which random forest partitions are crossed in order to reach it from the original instance. We demonstrate this method on both the MNIST hand-drawn digit dataset and the German credit dataset, finding that it generates explanations that are sparser and more useful than Shapley values.
Causal Masking on Spatial Data: An Information-Theoretic Case for Learning Spatial Datasets with Unimodal Language Models
Junkin, Jared, Nathanson, Samuel
Language models are traditionally designed around causal masking. In domains with spatial or relational structure, causal masking is often viewed as inappropriate, and sequential linearizations are instead used. Yet the question of whether it is viable to accept the information loss introduced by causal masking on nonsequential data has received little direct study, in part because few domains offer both spatial and sequential representations of the same dataset. In this work, we investigate this issue in the domain of chess, which naturally supports both representations. We train language models with bidirectional and causal self-attention mechanisms on both spatial (board-based) and sequential (move-based) data. Our results show that models trained on spatial board states - \textit{even with causal masking} - consistently achieve stronger playing strength than models trained on sequential data. While our experiments are conducted on chess, our results are methodological and may have broader implications: applying causal masking to spatial data is a viable procedure for training unimodal LLMs on spatial data, and in some domains is even preferable to sequentialization.
Robust fuzzy clustering for high-dimensional multivariate time series with outlier detection
Ma, Ziling, Lรณpez-Oriona, รngel, Ombao, Hernando, Sun, Ying
Fuzzy clustering provides a natural framework for modeling partial memberships, particularly important in multivariate time series (MTS) where state boundaries are often ambiguous. For example, in EEG monitoring of driver alertness, neural activity evolves along a continuum (from unconscious to fully alert, with many intermediate levels of drowsiness) so crisp labels are unrealistic and partial memberships are essential. However, most existing algorithms are developed for static, low-dimensional data and struggle with temporal dependence, unequal sequence lengths, high dimensionality, and contamination by noise or artifacts. To address these challenges, we introduce RFCPCA, a robust fuzzy subspace-clustering method explicitly tailored to MTS that, to the best of our knowledge, is the first of its kind to simultaneously: (i) learn membership-informed subspaces, (ii) accommodate unequal lengths and moderately high dimensions, (iii) achieve robustness through trimming, exponential reweighting, and a dedicated noise cluster, and (iv) automatically select all required hyperparameters. These components enable RFCPCA to capture latent temporal structure, provide calibrated membership uncertainty, and flag series-level outliers while remaining stable under contamination. On driver drowsiness EEG, RFCPCA improves clustering accuracy over related methods and yields a more reliable characterization of uncertainty and outlier structure in MTS.
Will AI mean the end of call centres?
Will AI mean the end of call centres? Ask ChatGPT whether AI will replace humans in the customer service industry, and it will offer a diplomatic answer, the summary of which is they will work side by side. Humans though, are not so optimistic. Last year, the chief executive of Indian technology firm Tata Consultancy Services, K Krithivasan, told the Financial Times that AI may soon mean that there is minimal need for call centres in Asia. Meanwhile, AI will autonomously resolve 80% of common customer service issues by 2029, predicts business and technology research firm Gartner.
'He lives for the goals' - robot Haaland returns from malfunction
'He lives for the goals' - robot Haaland returns from malfunction Is Erling Haaland a big fan of Peter Crouch - or is he actually programmed like a robot? That may - or not be - a question posed after Manchester City's impressive Premier League victory over in-form Bournemouth on Sunday. The Norway striker malfunctioned for only the second time this season when he failed to score in last weekend's loss at Aston Villa, but he was back to being a goal machine with a ruthlessly efficient first-half double against the Cherries. If he is hiding any nuts and bolts under those blonde locks of his, Haaland did prove he was still human by missing a couple of chances to complete his hat-trick. But his scary statistics this season have left many in awe of the 25-year-old's prowess in front of goal, prompting a robot dance to mark his opener in the win that took his side up to second place.
Slate Crossword: Where the Sharks and Jets Sometimes Duke It Out? (Three Letters)
Please enable Javascript in your browser to view Slate interactives. Today's puzzle is a 15x15 grid. Read about it in Slate: A look inside Midjourney, the mind-breaking A.I. tool where users create the sacred (Pope Francis in a puffer) and the profane (you don't want to know). Get Slate Games in your inbox every weekday. You can manage your newsletter subscriptions at any time.
What do you see? 12 extreme close-ups bring 'hidden science' to life
What do you see? 12 extreme close-ups bring'hidden science' to life New photography book encourages us to look for science everywhere. Over many years of geological time, opal is slowly formed as many small spheres of silica (what glass is made of) self-assemble into perfectly ordered layers. Breakthroughs, discoveries, and DIY tips sent every weekday. Organized into five thematic sections, the book turns learning about science into a guessing game. Detailed photographs like the ones below taken by MIT researcher and science photographer Felice Frankel challenge readers to deduce the underlying chemical, natural, or physical processes at play.
Hamas rejects US accusation it looted aid trucks in Gaza
Why did Israel launch air strikes on Gaza? What life is like in Gaza's crowded tents How is Israel using PR firms to frame its war? Will the US plan for Gaza fail? Hamas has denied accusations by the US Central Command (CENTCOM) that the Palestinian group looted aid trucks in the Gaza Strip. CENTCOM had published drone footage that allegedly showed an aid truck being looted in the enclave.