Education
Rerouting Connection: Hybrid Computer Vision Analysis Reveals Visual Similarity Between Indus and Tibetan-Yi Corridor Writing Systems
This thesis employs a hybrid CNN-Transformer architecture, alongside a detailed anthropological framework, to investigate potential historical connections between the visual morphology of the Indus Valley script and pictographic systems of the Tibetan-Yi Corridor. Through an ensemble methodology of three target scripts across 15 independently trained models, we demonstrate that Tibetan-Yi Corridor scripts exhibit approximately six-fold higher visual similarity to the Indus script (0.635) than to the Bronze Age Proto-Cuneiform (0.102) or Proto-Elamite (0.078). Contrary to expectations, when measured through direct script-to-script embedding comparisons, the Indus script maps closer to Tibetan-Yi Corridor scripts with a mean cosine similarity of 0.930 (CI: [0.917, 0.942]) than to contemporaneous West Asian signaries, which recorded mean similarities of 0.887 (CI: [0.863, 0.911]) and 0.855 (CI: [0.818, 0.891]). Across dimensionality reduction and clustering methods, the Indus script consistently clusters closest to Tibetan-Yi Corridor scripts. These computational findings align with observed pictorial parallels in numeral systems, gender markers, and iconographic elements. Archaeological evidence of contact networks along the ancient Shu-Shendu road, coinciding with the Indus Civilization's decline, provides a plausible transmission pathway. While alternate explanations cannot be ruled out, the specificity and consistency of similarities suggest more complex cultural transmission networks between South and East Asia than previously recognized.
Representation Learning by Ranking across multiple tasks
In recent years, representation learning has become the research focus of the machine learning community. Large-scale neural networks are a crucial step toward achieving general intelligence, with their success largely attributed to their ability to learn abstract representations of data. Several learning fields are actively discussing how to learn representations, yet there is a lack of a unified perspective. We convert the representation learning problem under different tasks into a ranking problem. By adopting the ranking problem as a unified perspective, representation learning tasks can be solved in a unified manner by optimizing the ranking loss. Experiments under various learning tasks, such as classification, retrieval, multi-label learning, and regression, prove the superiority of the representation learning by ranking framework. Furthermore, experiments under self-supervised learning tasks demonstrate the significant advantage of the ranking framework in processing unsupervised training data, with data augmentation techniques further enhancing its performance.
The Gen Z Lifestyle Subsidy
Finals season looks different this year. Across college campuses, students are slogging their way through exams with all-nighters and lots of caffeine, just as they always have. Through the end of May, OpenAI is offering students two months of free access to ChatGPT Plus, which normally costs 20 a month. It's a compelling deal for students who want help cramming--or cheating--their way through finals: Rather than firing up the free version of ChatGPT to outsource essay writing or work through a practice chemistry exam, students are now able to access the company's most advanced models, as well as its "deep research" tool, which can quickly synthesize hundreds of digital sources into analytical reports. The OpenAI deal is just one of many such AI promotions going around campuses.
Learn how to boss around AI bots before they become your boss
But AI is a tool; like any tool, it is only as good as the person wielding it. Now's the time to get the upper hand on AI and learn how to use tools like ChatGPT and automation platforms to work for you. The ChatGPT & Automation E-Degree from Eduonix Learning Solutions gives you the knowledge to stay on top for just 29.99 (MSRP 790) The course includes 12 modules and 25 hours of content you can move through at your own pace, and they never expire. You'll learn how to automate workflows, streamline repetitive tasks, and get AI to handle the boring stuff while you take credit for the results. It also dives into prompt engineering, real-world use cases, and customizing ChatGPT to fit your job, industry, or hustle.
Teens are now using AI chatbots to create and spread nude images of classmates, alarming education experts
A troubling trend has emerged in schools across the United States, with young students falling victim to the increasing use of artificial intelligence (AI)-powered "nudify" apps that have the power to create fake pornography of classmates. "Nudify" is an umbrella term referring to a plethora of widely available apps and websites that allow users to alter photos of full-dressed individuals and virtually undress them. Some apps can create nude images with just a headshot of the victim. Don Austin, the superintendent of the Palo Alto Unified School District, told Fox News Digital that this type of online harassment can be more relentless compared to traditional in-person bullying. "It used to be that a bully had to come over and push you. Palo Alto is not a community where people are going to come push anybody into a locker. But it's not immune from online bullying," Austin said.
Deep learning with missing data
Ma, Tianyi, Wang, Tengyao, Samworth, Richard J.
In the context of multivariate nonparametric regression with missing covariates, we propose Pattern Embedded Neural Networks (PENNs), which can be applied in conjunction with any existing imputation technique. In addition to a neural network trained on the imputed data, PENNs pass the vectors of observation indicators through a second neural network to provide a compact representation. The outputs are then combined in a third neural network to produce final predictions. Our main theoretical result exploits an assumption that the observation patterns can be partitioned into cells on which the Bayes regression function behaves similarly, and belongs to a compositional H\"older class. It provides a finite-sample excess risk bound that holds for an arbitrary missingness mechanism, and in combination with a complementary minimax lower bound, demonstrates that our PENN estimator attains in typical cases the minimax rate of convergence as if the cells of the partition were known in advance, up to a poly-logarithmic factor in the sample size. Numerical experiments on simulated, semi-synthetic and real data confirm that the PENN estimator consistently improves, often dramatically, on standard neural networks without pattern embedding. Code to reproduce our experiments, as well as a tutorial on how to apply our method, is publicly available.
Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models
Yang, Junjie, Song, Junhao, Han, Xudong, Bi, Ziqian, Wang, Tianyang, Liang, Chia Xin, Song, Xinyuan, Zhang, Yichao, Niu, Qian, Peng, Benji, Chen, Keyu, Liu, Ming
Knowledge distillation (KD) is a technique for transferring knowledge from complex teacher models to simpler student models, significantly enhancing model efficiency and accuracy. It has demonstrated substantial advancements in various applications including image classification, object detection, language modeling, text classification, and sentiment analysis. Recent innovations in KD methods, such as attention-based approaches, block-wise logit distillation, and decoupling distillation, have notably improved student model performance. These techniques focus on stimulus complexity, attention mechanisms, and global information capture to optimize knowledge transfer. In addition, KD has proven effective in compressing large language models while preserving accuracy, reducing computational overhead, and improving inference speed. This survey synthesizes the latest literature, highlighting key findings, contributions, and future directions in knowledge distillation to provide insights for researchers and practitioners on its evolving role in artificial intelligence and machine learning.
Large Language Models Will Change The Way Children Think About Technology And Impact Every Interaction Paradigm
It is a call for education and research in this space so that we can harness this irresistible force for more good than harm, and provides some early themes for designers to consider. We firstly discuss where and how LLMs have been used in school educational settings, and then explore the new opportunities that recently released models offer. A small-scale investigation reveals potentially large impacts on how children learn, and we highlight key things that we as a community need to be aware of. 2 A SIMPLE GUIDE TO LARGE LANGUAGE MODELS Large Language Models -- think ChatGPT, Gemini, GPT-3, CoPilot -- are immense deep learning neural networks with exceptional numbers of parameters, which are trained to pre dict sequences of words, having been trained on most of the contents of the Internet. If I asked you to complete the sentence Twinkle, twinkle, little star, how I wonder what you ..... it is quite likely that, if you have been brought up in a Wester n culture, you will recognise the nursery rhyme and complete the line with .....are LLMs do this, but on a massive scale. As the LLM has processed m uch of what has ever been written, it has ingested a large number of sequences of words, and compresses them to c reate an internal representation. An LLM can be seen as the JPEG of the web -- it is a lossy compressed version of the internet.
Simulating Before Planning: Constructing Intrinsic User World Model for User-Tailored Dialogue Policy Planning
He, Tao, Liao, Lizi, Liu, Ming, Qin, Bing
Recent advancements in dialogue policy planning have emphasized optimizing system agent policies to achieve predefined goals, focusing on strategy design, trajectory acquisition, and efficient training paradigms. However, these approaches often overlook the critical role of user characteristics, which are essential in real-world scenarios like conversational search and recommendation, where interactions must adapt to individual user traits such as personality, preferences, and goals. To address this gap, we first conduct a comprehensive study utilizing task-specific user personas to systematically assess dialogue policy planning under diverse user behaviors. By leveraging realistic user profiles for different tasks, our study reveals significant limitations in existing approaches, highlighting the need for user-tailored dialogue policy planning. Building on this foundation, we present the User-Tailored Dialogue Policy Planning (UDP) framework, which incorporates an Intrinsic User World Model to model user traits and feedback. UDP operates in three stages: (1) User Persona Portraying, using a diffusion model to dynamically infer user profiles; (2) User Feedback Anticipating, leveraging a Brownian Bridge-inspired anticipator to predict user reactions; and (3) User-Tailored Policy Planning, integrating these insights to optimize response strategies. To ensure robust performance, we further propose an active learning approach that prioritizes challenging user personas during training. Comprehensive experiments on benchmarks, including collaborative and non-collaborative settings, demonstrate the effectiveness of UDP in learning user-specific dialogue strategies. Results validate the protocol's utility and highlight UDP's robustness, adaptability, and potential to advance user-centric dialogue systems.
Bayesian continual learning and forgetting in neural networks
Bonnet, Djohan, Cottart, Kellian, Hirtzlin, Tifenn, Januel, Tarcisius, Dalgaty, Thomas, Vianello, Elisa, Querlioz, Damien
Biological synapses effortlessly balance memory retention and flexibility, yet artificial neural networks still struggle with the extremes of catastrophic forgetting and catastrophic remembering. Here, we introduce Metaplasticity from Synaptic Uncertainty (MESU), a Bayesian framework that updates network parameters according their uncertainty. This approach allows a principled combination of learning and forgetting that ensures that critical knowledge is preserved while unused or outdated information is gradually released. Unlike standard Bayesian approaches -- which risk becoming overly constrained, and popular continual-learning methods that rely on explicit task boundaries, MESU seamlessly adapts to streaming data. It further provides reliable epistemic uncertainty estimates, allowing out-of-distribution detection, the only computational cost being to sample the weights multiple times to provide proper output statistics. Experiments on image-classification benchmarks demonstrate that MESU mitigates catastrophic forgetting, while maintaining plasticity for new tasks. When training 200 sequential permuted MNIST tasks, MESU outperforms established continual learning techniques in terms of accuracy, capability to learn additional tasks, and out-of-distribution data detection. Additionally, due to its non-reliance on task boundaries, MESU outperforms conventional learning techniques on the incremental training of CIFAR-100 tasks consistently in a wide range of scenarios. Our results unify ideas from metaplasticity, Bayesian inference, and Hessian-based regularization, offering a biologically-inspired pathway to robust, perpetual learning.