Large Language Model
A vibe coding learning design to enhance EFL students' talking to, through, and about AI
Woo, David James, Guo, Kai, Yu, Yangyang
This innovative practice article reports on the piloting of vibe coding (using natural language to create software applications with AI) for English as a Foreign Language (EFL) education. We developed a human-AI meta-languaging framework with three dimensions: talking to AI (prompt engineering), talking through AI (negotiating authorship), and talking about AI (mental models of AI). Using backward design principles, we created a four-hour workshop where two students designed applications addressing authentic EFL writing challenges. We adopted a case study methodology, collecting data from worksheets and video recordings, think-aloud protocols, screen recordings, and AI-generated images. Contrasting cases showed one student successfully vibe coding a functional application cohering to her intended design, while another encountered technical difficulties with major gaps between intended design and actual functionality. Analysis reveals differences in students' prompt engineering approaches, suggesting different AI mental models and tensions in attributing authorship. We argue that AI functions as a beneficial languaging machine, and that differences in how students talk to, through, and about AI explain vibe coding outcome variations. Findings indicate that effective vibe coding instruction requires explicit meta-languaging scaffolding, teaching structured prompt engineering, facilitating critical authorship discussions, and developing vocabulary for articulating AI mental models.
PerFairX: Is There a Balance Between Fairness and Personality in Large Language Model Recommendations?
The integration of Large Language Models (LLMs) into recommender systems has enabled zero-shot, personality-based personalization through prompt-based interactions, offering a new paradigm for user-centric recommendations. However, incorporating user personality traits via the OCEAN model highlights a critical tension between achieving psychological alignment and ensuring demographic fairness. To address this, we propose PerFairX, a unified evaluation framework designed to quantify the trade-offs between personalization and demographic equity in LLM-generated recommendations. Using neutral and personality-sensitive prompts across diverse user profiles, we benchmark two state-of-the-art LLMs, ChatGPT and DeepSeek, on movie (MovieLens 10M) and music (Last.fm 360K) datasets. Our results reveal that personality-aware prompting significantly improves alignment with individual traits but can exacerbate fairness disparities across demographic groups. Specifically, DeepSeek achieves stronger psychological fit but exhibits higher sensitivity to prompt variations, while ChatGPT delivers stable yet less personalized outputs. PerFairX provides a principled benchmark to guide the development of LLM-based recommender systems that are both equitable and psychologically informed, contributing to the creation of inclusive, user-centric AI applications in continual learning contexts.
MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
Yang, Garry, Chen, Zizhe, Wong, Man Hon, Lei, Haoyu, Chen, Yongqiang, Li, Zhenguo, Zhou, Kaiwen, Cheng, James
Large Video Models (LVMs) build on the semantic capabilities of Large Language Models (LLMs) and vision modules by integrating temporal information to better understand dynamic video content. Despite their progress, LVMs are prone to hallucinations-producing inaccurate or irrelevant descriptions. Current benchmarks for video hallucination depend heavily on manual categorization of video content, neglecting the perception-based processes through which humans naturally interpret videos. We introduce MESH, a benchmark designed to evaluate hallucinations in LVMs systematically. MESH uses a Question-Answering framework with binary and multi-choice formats incorporating target and trap instances. It follows a bottom-up approach, evaluating basic objects, coarse-to-fine subject features, and subject-action pairs, aligning with human video understanding. We demonstrate that MESH offers an effective and comprehensive approach for identifying hallucinations in videos. Our evaluations show that while LVMs excel at recognizing basic objects and features, their susceptibility to hallucinations increases markedly when handling fine details or aligning multiple actions involving various subjects in longer videos.
Towards Generalized Routing: Model and Agent Orchestration for Adaptive and Efficient Inference
Guo, Xiyu, Wang, Shan, Ji, Chunfang, Zhao, Xuefeng, Xi, Wenhao, Liu, Yaoyao, Li, Qinglan, Deng, Chao, Feng, Junlan
The rapid advancement of large language models (LLMs) and domain-specific AI agents has greatly expanded the ecosystem of AI-powered services. User queries, however, are highly diverse and often span multiple domains and task types, resulting in a complex and heterogeneous landscape. This diversity presents a fundamental routing challenge: how to accurately direct each query to an appropriate execution unit while optimizing both performance and efficiency. To address this, we propose MoMA (Mixture of Models and Agents), a generalized routing framework that integrates both LLM and agent-based routing. Built upon a deep understanding of model and agent capabilities, MoMA effectively handles diverse queries through precise intent recognition and adaptive routing strategies, achieving an optimal balance between efficiency and cost. Specifically, we construct a detailed training dataset to profile the capabilities of various LLMs under different routing model structures, identifying the most suitable tasks for each LLM. During inference, queries are dynamically routed to the LLM with the best cost-performance efficiency. We also introduce an efficient agent selection strategy based on a context-aware state machine and dynamic masking. Experimental results demonstrate that the MoMA router offers superior cost-efficiency and scalability compared to existing approaches.
Anthropic's Claude AI chatbot can now create and edit Office files
When you purchase through links in our articles, we may earn a small commission. Anthropic's Claude AI chatbot can now create and edit Office files Claude AI is now more of an active collaborator, says Anthropic. According to an announcement post, Anthropic has launched a new feature in Claude that allows you to create and edit files directly in the AI's chat--including Word documents, Excel spreadsheets, PowerPoint presentations, and PDFs. Previously, only basic file support was offered. Through a private computing environment, Claude can now write code and run programs to generate files and analyses.
Partnering with generative AI in the finance function
CFOs are experimenting with AI use cases to free up capacity for business-critical work. Generative AI has the potential to transform the finance function. By taking on some of the more mundane tasks that can occupy a lot of time, generative AI tools can help free up capacity for more high-value strategic work. For chief financial officers, this could mean spending more time and energy on proactively advising the business on financial strategy as organizations around the world continue to weather ongoing geopolitical and financial uncertainty. CFOs can use large language models (LLMs) and generative AI tools to support everyday tasks like generating quarterly reports, communicating with investors, and formulating strategic summaries, says Andrew W. Lo, Charles E. and Susan T. Harris professor and director of the Laboratory for Financial Engineering at the MIT Sloan School of Management. "LLMs can't replace the CFO by any means, but they can take a lot of the drudgery out of the role by providing first drafts of documents that summarize key issues and outline strategic priorities."
The Download: Trump's impact on science, and meet our climate and energy honorees
The Download: Trump's impact on science, and meet our climate and energy honorees How Trump's policies are affecting early-career scientists--in their own words Every year MIT Technology Review celebrates accomplished young scientists, entrepreneurs, and inventors from around the world in our Innovators Under 35 list. We've just published the 2025 edition . This year, though, the context is different: The US scientific community is under attack. Since Donald Trump took office in January, his administration has fired top government scientists, targeted universities and academia, and made substantial funding cuts to the country's science and technology infrastructure. We asked our six most recent cohorts about both positive and negative impacts of the administration's new policies. Their responses provide a glimpse into the complexities of building labs, companies, and careers in today's political climate.
How thousands of 'overworked, underpaid' humans train Google's AI to seem smart
AI models are trained on vast swathes of data from every corner of the internet, by humans. AI models are trained on vast swathes of data from every corner of the internet, by humans. How thousands of'overworked, underpaid' humans train Google's AI to seem smart In the spring of 2024, when Rachael Sawyer, a technical writer from Texas, received a LinkedIn message from a recruiter hiring for a vague title of writing analyst, she assumed it would be similar to her previous gigs of content creation. On her first day a week later, however, her expectations went bust. Instead of writing words herself, Sawyer's job was to rate and moderate the content created by artificial intelligence. The job initially involved a mix of parsing through meeting notes and chats summarized by Google's Gemini, and, in some cases, reviewing short films made by the AI.
Apertus: a fully open, transparent, multilingual language model
In July, EPFL, ETH Zurich, and the Swiss National Supercomputing Centre (CSCS) announced their joint initiative to build a large language model (LLM) . Now, this model is available and serves as a building block for developers and organisations for future applications such as chatbots, translation systems, or educational tools. The model is named Apertus - Latin for "open" - highlighting its distinctive feature: the entire development process, including its architecture, model weights, and training data and recipes, is openly accessible and fully documented. AI researchers, professionals, and experienced enthusiasts can either access the model through the strategic partner Swisscom or download it from Hugging Face - a platform for AI models and applications - and deploy it for their own projects. Apertus is freely available in two sizes - featuring 8 billion and 70 billion parameters, the smaller model being more appropriate for individual usage.
TopResume Free Review, Discounts & Packages for September 2025
Discover ways to save at TopResume, including their free review service and 4-week Career Services Platform trial. All products featured on WIRED are independently selected by our editors. However, we may receive compensation from retailers and/or from purchases of products through these links. AI is making it harder to find a job. AI-driven Application Tracking Systems (ATS) can dump your resume before a recruiter has ever seen it, even if you have all of your qualifications clearly spelled out.