Goto

Collaborating Authors

 Cognitive Science


Remember Jibo? Its Successor Is a Wearable That Turns Your Life Into AI Slop

WIRED

With "blessings" from the original Jibo founders, iKairos is a wearable or desk-mounted "AI journal" that turns your family moments into AI images. Jibo was a cute, social robot that sat in one place in your home. Built in 2014, Jibo was meant to be a robot people brought into their lives before the smart-home industry had even really taken off. It could not interact with objects or move from its spot, but it could wiggle and talk in a way that felt endearing. It was adorable, but not competent or in demand enough to last--it was ahead of its time and died too soon .


The Download: Claude's inner workings, and the future of world models

MIT Technology Review

Plus: New York has become the first state to enact a data center moratorium. When Anthropic announced last week that it had found a new window into its models' "internal thoughts" as they reason through answers, there was one colleague I had to talk to: senior editor Will Douglas Heaven. Aside from having a PhD in computer science, Will has spent a lot of time digging into what we can say about how AI models work. I spoke with him about what we should take from Anthropic's new (and typically quirky) research. Here's what he had to say . How will AI understand the real world?


The Download: a donor conception cap and world models for AI

MIT Technology Review

Plus: Apple has sued OpenAI for allegedly stealing trade secrets. Ties van der Meer doesn't know how many siblings he has. The 47-year-old was conceived at a private fertility clinic using sperm from an anonymous donor. He eventually tracked down one sibling, but he may have others he'll never find. Other donor-conceived people have found they have tens or even hundreds of them. "It does make you feel a bit mass-produced," said one who discovered they had 25 half-siblings.


ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding

Neural Information Processing Systems

Charts are high-density visualization carriers for complex data, serving as a crucial medium for information extraction and analysis. Automated chart understanding poses significant challenges to existing multimodal large language models (MLLMs) due to the need for precise and complex visual reasoning. Current step-by-step reasoning models primarily focus on text-based logical reasoning for chart understanding. However, they struggle to refine or correct their reasoning when errors stem from flawed visual understanding, as they lack the ability to leverage multimodal interaction for deeper comprehension. Inspired by human cognitive behavior, we propose ChartSketcher, a multimodal feedback-driven step-by-step reasoning method designed to address these limitations. ChartSketcher is a chart understanding model that employs Sketch-CoT, enabling MLLMs to annotate intermediate reasoning steps directly onto charts using a programmatic sketching library, iteratively feeding these visual annotations back into the reasoning process. This mechanism enables the model to visually ground its reasoning and refine its understanding over multiple steps. We employ a two-stage training strategy: a cold start phase to learn sketch-based reasoning patterns, followed by off-policy reinforcement learning to enhance reflection and generalization. Experiments demonstrate that ChartSketcher achieves promising performance on chart understanding benchmarks and general vision tasks, providing an interactive and interpretable approach to chart comprehension.


StateSpaceDiffuser: Bringing Long Context to Diffusion World Models

Neural Information Processing Systems

World models have recently gained prominence for action-conditioned visual prediction in complex environments. However, relying on only a few recent observations causes them to lose long-term context. Consequently, within a few steps, the generated scenes drift from what was previously observed, undermining temporal coherence. This limitation, common in state-of-the-art world models, which are diffusion-based, stems from the lack of a lasting environment state. To address this problem, we introduce StateSpaceDiffuser, where a diffusion model is enabled to perform long-context tasks by integrating features from a state-space model, representing the entire interaction history.


Neuroscience can't tell us the way to govern people's brains

New Scientist

Neuroscience can't tell us the way to govern people's brains From the age of legal adulthood to the concept of profound autism, policy-makers are turning to neuroscience to help shape laws and policies, but the science simply isn't ready Decisions are often made via a subconscious muddling through, due to the brain's desire to minimise energy use . It is perhaps why we value neat categorisations of someone's brain state, despite these being flawed. Take the age at which you become an adult. Around the world, legal adulthood varies from 16 to 21. This difference matters, as we rightly have different expectations for children versus adults.


Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties

Neural Information Processing Systems

Recent large-scale reasoning models have achieved state-of-the-art performance on challenging mathematical benchmarks, yet the internal mechanisms underlying their success remain poorly understood. In this work, we introduce the notion of a reasoning graph, extracted by clustering hidden-state representations at each reasoning step, and systematically analyze three key graph-theoretic properties: cyclicity, diameter, and small-world index, across multiple tasks (GSM8K, MATH500, AIME 2024). Our findings reveal that distilled reasoning models (e.g., DeepSeekR1-Distill-Qwen-32B) exhibit significantly more recurrent cycles (about 5 per sample), substantially larger graph diameters, and pronounced small-world characteristics (about 6x) compared to their base counterparts. Notably, these structural advantages grow with task difficulty and model capacity, with cycle detection peaking at the 14B scale and exploration diameter maximized in the 32B variant, correlating positively with accuracy. Furthermore, we show that supervised fine-tuning on an improved dataset systematically expands reasoning graph diameters in tandem with performance gains, offering concrete guidelines for dataset design aimed at boosting reasoning capabilities.


In Silico Mapping of Visual Categorical Selectivity Across the Whole Brain

Neural Information Processing Systems

A fine-grained account of functional selectivity in the cortex is essential for understanding how visual information is processed and represented in the brain. Classical studies using designed experiments have identified multiple category-selective regions; however, these approaches rely on preconceived hypotheses about categories. Subsequent data-driven discovery methods have sought to address this limitation but are often limited by simple, typically linear encoding models. We propose an in silico approach for data-driven discovery of novel category-selectivity hypotheses based on an encoder-decoder transformer model. The architecture incorporates a brain-region to image-feature cross-attention mechanism, enabling nonlinear mappings between high-dimensional deep network features and semantic patterns encoded in the brain activity. We further introduce a method to characterize the selectivity of individual parcels by leveraging diffusion-based image generative models and large-scale datasets to synthesize and select images that maximally activate each parcel. Our approach reveals regions with complex, compositional selectivity involving diverse semantic concepts, which we validate in silico both within and across subjects. Using a brain encoder as a "digital twin" offers a powerful, data-driven framework for generating and testing hypotheses about visual selectivity in the human brain--hypotheses that can guide future fMRI experiments.


CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists ' Diagnostic Logic

Neural Information Processing Systems

Recent advances in computational pathology have led to the emergence of numerous foundation models. These models typically rely on general-purpose encoders with multi-instance learning for whole slide image (WSI) classification or apply multimodal approaches to generate reports directly from images. However, these models cannot emulate the diagnostic approach of pathologists, who systematically examine slides at low magnification to obtain an overview before progressively zooming in on suspicious regions to formulate comprehensive diagnoses.


3 People Have Gotten Cancer-Detecting Implants in Their Brains

WIRED

The startup Coherence Neuro is now testing a brain-computer interface that could one day use electrical stimulation to prevent tumors from growing. A San Francisco startup with ties to Elon Musk's Neuralink has started testing its brain implant to detect and treat cancer in humans. Coherence Neuro says it temporarily placed its coin-sized implant in the brains of three people undergoing surgery to have brain tumors removed at the Royal Melbourne Hospital in Australia. The implant was in place for roughly 30 minutes before being removed, providing an important safety check before the device can be implanted long-term in patients with brain cancer. Known as a brain-computer interface, the Coherence Neuro device is designed to sense the unique electrical signals of tumors and deliver mild electrical stimulation to prevent their growth.