scene template
Spontaneous High-Order Generalization in Neural Theory-of-Mind Networks
Theory-of-Mind (ToM) is a core human cognitive capacity for attributing mental states to self and others. Wimmer and Perner demonstrated that humans progress from first- to higher-order ToM within a short span, completing this development before formal education or advanced skill acquisition. In contrast, neural networks represented by autoregressive language models progress from first- to higher-order ToM only alongside gains in advanced skills like reasoning, leaving open whether their trajectory can unfold independently, as in humans. In this research, we provided evidence that neural networks could spontaneously generalize from first- to higher-order ToM without relying on advanced skills. We introduced a neural Theory-of-Mind network (ToMNN) that simulated a minimal cognitive system, acquiring only first-order ToM competence. Evaluations of its second- and third-order ToM abilities showed accuracies well above chance. Also, ToMNN exhibited a sharper decline when generalizing from first- to second-order ToM than from second- to higher orders, and its accuracy decreased with greater task complexity. These perceived difficulty patterns were aligned with human cognitive expectations. Furthermore, the universality of results was confirmed across different parameter scales. Our findings illuminate machine ToM generalization patterns and offer a foundation for developing more human-like cognitive systems.
Flow-Guided Video Inpainting with Scene Templates
Lao, Dong, Zhu, Peihao, Wonka, Peter, Sundaramoorthi, Ganesh
We consider the problem of filling in missing spatio-temporal regions of a video. We provide a novel flow-based solution by introducing a generative model of images in relation to the scene (without missing regions) and mappings from the scene to images. We use the model to jointly infer the scene template, a 2D representation of the scene, and the mappings. This ensures consistency of the frame-to-frame flows generated to the underlying scene, reducing geometric distortions in flow based inpainting. The template is mapped to the missing regions in the video by a new L2-L1 interpolation scheme, creating crisp inpaintings and reducing common blur and distortion artifacts. We show on two benchmark datasets that our approach out-performs state-of-the-art quantitatively and in user studies.
SketchyScene: Richly-Annotated Scene Sketches
Zou, Changqing, Yu, Qian, Du, Ruofei, Mo, Haoran, Song, Yi-Zhe, Xiang, Tao, Gao, Chengying, Chen, Baoquan, Zhang, Hao
We contribute the first large-scale dataset of scene sketches, SketchyScene, with the goal of advancing research on sketch understanding at both the object and scene level. The dataset is created through a novel and carefully designed crowdsourcing pipeline, enabling users to efficiently generate large quantities of realistic and diverse scene sketches. SketchyScene contains more than 29,000 scene-level sketches, 7,000+ pairs of scene templates and photos, and 11,000+ object sketches. All objects in the scene sketches have ground-truth semantic and instance masks. The dataset is also highly scalable and extensible, easily allowing augmenting and/or changing scene composition. We demonstrate the potential impact of SketchyScene by training new computational models for semantic segmentation of scene sketches and showing how the new dataset enables several applications including image retrieval, sketch colorization, editing, and captioning, etc. The dataset and code can be found at https://github.com/SketchyScene/SketchyScene.
Video games where people matter? The strange future of emotional AI - IBM for Games
Video games where people matter? If you're a video game fan of a certain age, you may remember Edge magazine's controversial review of the bloody sci-fi shooting game, Doom. Perhaps you enjoyed a good laugh, as many first-person shooter fans have, at the writer's much-mocked assertion: "if only you could talk to these creatures, then perhaps you could try and make friends with them, form alliances … Now that would be interesting." Of course, we all know what happened. There would be no room in the Doom series, nor any subsequent first-person blast-'em-up, for such socio-psychological niceties. Instead, we enjoyed 20 years of shooting, bludgeoning and stabbing, the ludicrous idea of diplomacy cast roughly aside. But during this era, something else was happening in game design, and in academic thinking around video games and artificial intelligence.
Video games where people matter? The strange future of emotional AI
If you're a video game fan of a certain age, you may remember Edge magazine's controversial review of the bloody sci-fi shooting game, Doom. Perhaps you enjoyed a good laugh, as many first-person shooter fans have, at the writer's much-mocked assertion: "if only you could talk to these creatures, then perhaps you could try and make friends with them, form alliances ... Now that would be interesting." Of course, we all know what happened. There would be no room in the Doom series, nor any subsequent first-person blast-'em-up, for such socio-psychological niceties. Instead, we enjoyed 20 years of shooting, bludgeoning and stabbing, the ludicrous idea of diplomacy cast roughly aside. But during this era, something else was happening in game design, and in academic thinking around video games and artificial intelligence. Buoyed by advances in AI research and aided by increasingly powerful computer processors, developers were beginning to think about the possibilities of non-player characters (NPCs) who could think and act in a more complex and human way – who could provide the emotional feedback that the Edge reviewer was thinking about.