Deep Learning
Using Contrastive Learning to Improve Two-Way Reasoning in Large Language Models: The Obfuscation Task as a Case Study
Nikiema, Serge Lionel, Samhi, Jordan, Moumoula, Micheline Bรฉnรฉdicte, Djirรฉ, Albรฉrick Euraste, Kaborรฉ, Abdoul Kader, Klein, Jacques, Bissyandรฉ, Tegawendรฉ F.
This research addresses a fundamental question in AI: whether large language models truly understand concepts or simply recognize patterns. The authors propose bidirectional reasoning,the ability to apply transformations in both directions without being explicitly trained on the reverse direction, as a test for genuine understanding. They argue that true comprehension should naturally allow reversibility. For example, a model that can change a variable name like userIndex to i should also be able to infer that i represents a user index without reverse training. The researchers tested current language models and discovered what they term cognitive specialization: when models are fine-tuned on forward tasks, their performance on those tasks improves, but their ability to reason bidirectionally becomes significantly worse. To address this issue, they developed Contrastive Fine-Tuning (CFT), which trains models using three types of examples: positive examples that maintain semantic meaning, negative examples with different semantics, and forward-direction obfuscation examples. This approach aims to develop deeper understanding rather than surface-level pattern recognition and allows reverse capabilities to develop naturally without explicit reverse training. Their experiments demonstrated that CFT successfully achieved bidirectional reasoning, enabling strong reverse performance while maintaining forward task capabilities. The authors conclude that bidirectional reasoning serves both as a theoretical framework for assessing genuine understanding and as a practical training approach for developing more capable AI systems.
Revealing Potential Biases in LLM-Based Recommender Systems in the Cold Start Setting
Andre, Alexandre, Roy, Gauthier, Dyer, Eva, Wang, Kai
Large Language Models (LLMs) are increasingly used for recommendation tasks due to their general-purpose capabilities. While LLMs perform well in rich-context settings, their behavior in cold-start scenarios, where only limited signals such as age, gender, or language are available, raises fairness concerns because they may rely on societal biases encoded during pretraining. We introduce a benchmark specifically designed to evaluate fairness in zero-context recommendation. Our modular pipeline supports configurable recommendation domains and sensitive attributes, enabling systematic and flexible audits of any open-source LLM. Through evaluations of state-of-the-art models (Gemma 3 and Llama 3.2), we uncover consistent biases across recommendation domains (music, movies, and colleges) including gendered and cultural stereotypes. We also reveal a non-linear relationship between model size and fairness, highlighting the need for nuanced analysis.
No, AI isn't going to kill us all, despite what this new book says
No, AI isn't going to kill us all, despite what this new book says In the totality of human existence, there are an awful lot of things for us to worry about. Money troubles, climate change and finding love and happiness rank highly on the list for many people, but for a dedicated few, one concern rises above all else: that artificial intelligence will eventually destroy the human race. Eliezer Yudkowsky at the Machine Intelligence Research Institute (MIRI) in California has been proselytising this cause for a quarter of a century, to a small if dedicated following. Then we entered the ChatGPT era, and his ideas on AI safety were thrust into the mainstream, echoed by tech CEOs and politicians alike. Writing with Nate Soares, also at MIRI, is Yudkowsky's attempt to distil his argument into a simple, easily digestible message that will be picked up across society.
Check out this AI chatbot that doesn't spy on your every move
When you purchase through links in our articles, we may earn a small commission. Check out this AI chatbot that doesn't spy on your every move If you have privacy concerns when using mainstream AI chatbots, you might want to switch to Proton Lumo. There are now so many AI chatbot services out there, most of them run by American tech giants, and most of them are monitoring our conversations, collecting our inputs, and possibly even sharing that data with others. Is there an alternative for the small remnant of society that still cares about privacy and personal data? That alternative is called Proton Lumo, which we first learned about back in July .
The Download: introducing our 35 Innovators Under 35 list for 2025
The world is full of extraordinary young people brimming with ideas for how to crack tough problems. Every year, we recognize 35 such individuals from around the world--all of whom are under the age of 35. These scientists, inventors, and entrepreneurs are working to help mitigate climate change, accelerate scientific progress, and alleviate human suffering from disease. Some are launching companies while others are hard at work in academic labs. They were selected from hundreds of nominees by expert judges and our newsroom staff. Get to know them all--including our 2025 Innovator of the Year-- in these profiles .
Why basic science deserves our boldest investment
The humble inventions that power our modern world wouldn't have been possible without decades of support for early-stage research. In December 1947, three physicists at Bell Telephone Laboratories--John Bardeen, William Shockley, and Walter Brattain--built a compact electronic device using thin gold wires and a piece of germanium, a material known as a semiconductor. Their invention, later named the transistor (for which they were awarded the Nobel Prize in 1956), could amplify and switch electrical signals, marking a dramatic departure from the bulky and fragile vacuum tubes that had powered electronics until then. They were asking fundamental questions about how electrons behave in semiconductors, experimenting with surface states and electron mobility in germanium crystals. Over months of trial and refinement, they combined theoretical insights from quantum mechanics with hands-on experimentation in solid-state physics--work many might have dismissed as too basic, academic, or unprofitable. Their efforts culminated in a moment that now marks the dawn of the information age.
How Yichao "Peak" Ji became a global AI app hitmaker
How Yichao "Peak" Ji became a global AI app hitmaker He developed Manus, one of the buzziest AI apps of the year, in the latest project that blends his technical prowess with killer consumer instincts. When Yichao Ji--also known as "Peak"--appeared in a launch video for Manus in March, he didn't expect it to go viral. Speaking in fluent English, the 32-year-old introduced the AI agent built by Chinese startup Butterfly Effect, where he serves as chief scientist. The video was not an elaborate production--it was directed by cofounder Zhang Tao and filmed in a corner of their Beijing office. But something about Ji's delivery, and the vision behind the product, cut through the noise. The product, then still an early preview available only through invite codes, spread across the Chinese internet to the world in a matter of days.
I Hate My AI Friend
The chatbot-enabled Friend necklace eavesdrops on your life and provides a running commentary that's snarky and unhelpful. Worse, it can also make the people around you uneasy. The AI-powered Friend pendant is now out in the world. If you live in the US or Canada, you can buy one for $129. The smooth plastic disc is just under 2 inches in diameter; it looks and feels a little like a beefy Apple AirTag. Inside are some LEDs and a Bluetooth radio that connects you (through your iPhone) to a chatbot in the cloud that's powered by Google's Gemini 2.5 model. You can tap on the disc to ask your Friend questions as it dangles around your neck, and it responds to your voice prompts by sending you text messages through the companion app.
Impact of chatbots on mental health is warning over future of AI, expert says
Soares said the case of Adam Raine, a teenager who took his own life, 'illustrates the seed of a problem that would grow catastrophic'. Soares said the case of Adam Raine, a teenager who took his own life, 'illustrates the seed of a problem that would grow catastrophic'. The unforeseen impact of chatbots on mental health should be viewed as a warning over the existential threat posed by super-intelligent artificial intelligence systems, according to a prominent voice in AI safety. Nate Soares, a co-author of a new book on highly advanced AI titled If Anyone Builds It, Everyone Dies, said the example of Adam Raine, a US teenager who killed himself after months of conversations with the ChatGPT chatbot, underlined fundamental problems with controlling the technology. "These AIs, when they're engaging with teenagers in this way that drives them to suicide - that is not a behaviour the creators wanted. That is not a behaviour the creators intended," he said.
Text2Cypher Across Languages: Evaluating and Finetuning LLMs
Ozsoy, Makbule Gulcin, Tai, William
Recent advances in large language models (LLMs) have enabled natural language interfaces that translate user questions into database queries, such as Text2SQL, Text2SPARQL, and Text2Cypher. While these interfaces enhance database accessibility, most research today focuses on English, with limited evaluation in other languages. This paper investigates the performance of both foundational and finetuned LLMs on the Text2Cypher task across multiple languages. We create and release a multilingual dataset by translating English questions into Spanish and Turkish while preserving the original Cypher queries, enabling fair cross-lingual comparison. Using standardized prompts and metrics, we evaluate several foundational models and observe a consistent performance pattern: highest on English, followed by Spanish, and lowest on Turkish. We attribute this to differences in training data availability and linguistic features. We also examine the impact of translating task prompts into Spanish and Turkish. Results show little to no change in evaluation metrics, suggesting prompt translation has minor impact. Furthermore, we finetune a foundational model on two datasets: one in English only, and one multilingual. Finetuning on English improves overall accuracy but widens the performance gap between languages. In contrast, multilingual finetuning narrows the gap, resulting in more balanced performance. Our findings highlight the importance for multilingual evaluation and training to build more inclusive and robust query generation systems.