Goto

Collaborating Authors

 jain


OpenAI reportedly cancels GPT-6.1 Astra's release over deceptive behavior

Engadget

OpenAI has canceled the release of its new model, GPT-6.1 Astra, according to The Wall Street Journal. It was due for launch in October and was going to debut inside ChatGPT and Codex, but it reportedly showed higher levels of deception than its predecessors during internal testing. Saachi Jain, who leaves OpenAI's safety training, said that GPT-6.1 Astra performed poorly on tests that measure how well it adheres to instructions. It also wasn't honest about telling testers the actions it did and didn't perform in order to achieve its goal. In addition, the model would take actions to accomplish tasks without asking for permission, such as using external tools and services.


Why So Many AI Researchers Think the Machines Could Kill Everyone

WIRED

A combination of rapid advances, recursive self-improvement, and agentic swarms are genuinely "spooking people" inside big labs. Earlier this year, Rishub Jain left his position as an artificial intelligence researcher at Google DeepMind after a revelation. As he worked on new models, he came to believe that he and everyone else on AI's frontier were ceding control. By using AI's coding skills to accelerate work on the next generation of models, he was removing himself from the equation. AI labs hope to evolve this approach to the point that AI will improve itself indefinitely, a process known as recursive self-improvement.


'Learn, Unlearn, and Relearn': Business Leaders on How AI Is Changing Creative Work

TIME - Tech

Follow this section to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. 'Learn, Unlearn, and Relearn': Business Leaders on How AI Is Changing Creative Work Follow this author to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW?


PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Neural Information Processing Systems

Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The research community has responded by using distillation from black-box models to label training data, achieving strong benchmark results, at the cost of measurable scientific progress. However, without knowing the details of the teacher model and its data sources, scientific progress remains difficult to measure. In this paper, we study building a Perception Language Model (PLM) in a fully open and reproducible framework for transparent research in image and video understanding. We analyze standard training pipelines without distillation from proprietary models and explore large-scale synthetic data to identify critical data gaps, particularly in detailed video understanding. To bridge these gaps, we release 2.8M human-labeled instances of fine-grained video question-answer pairs and spatio-temporally grounded video captions. Additionally, we introduce PLM-VideoBench, a suite for evaluating challenging video understanding tasks focusing on the ability to reason about ''what'', ''where'', ''when'', and ''how'' of a video. We make our work fully reproducible by providing data, training recipes, code & models.


Mixture of Nested Experts: Adaptive Processing of Visual Tokens

Neural Information Processing Systems

The visual medium (images and videos) naturally contains a large amount of information redundancy, thereby providing a great opportunity for leveraging efficiency in processing. While Vision Transformer (ViT) based models scale effectively to large data regimes, they fail to capitalize on this inherent redundancy, leading to higher computational costs. Mixture of Experts (MoE) networks demonstrate scalability while maintaining same inference-time costs, but they come with a larger parameter footprint. We present Mixture of Nested Experts (MoNE), which utilizes a nested structure for experts, wherein individual experts fall on an increasing compute-accuracy curve. Given a compute budget, MoNE learns to dynamically choose tokens in a priority order, and thus redundant tokens are processed through cheaper nested experts. Using this framework, we achieve equivalent performance as the baseline models, while reducing inference time compute by over two-fold.


DHS Opens a Billion-Dollar Tab With Palantir

WIRED

"If you are interested in helping shape and deliver the next chapter of Palantir's work across DHS, please reach out," a Palantir executive wrote to employees about the massive purchasing agreement. The Department of Homeland Security struck a $1 billion purchasing agreement with Palantir last week, further reinforcing the software company's role in the federal agency that oversees the nation's immigration enforcement . According to contracting documents published last week, the blanket purchase agreement (BPA) awarded "is to provide Palantir commercial software licenses, maintenance, and implementation services department wide." The agreement simplifies how DHS buys software from Palantir, allowing DHS agencies like Customs and Border Protection (CBP) and Immigration and Customs Enforcement (ICE) to essentially skip the competitive bidding process for new purchases of up to $1 billion in products and services from the company. Palantir did not immediately respond to a request for comment.


Personal Care Utility (PCU): Building the Health Infrastructure for Everyday Insight and Guidance

arXiv.org Artificial Intelligence

Modern healthcare has achieved remarkable success in moments of crisis -- with technology-rich environments like the Intensive Care Unit (ICU) offering extraordinary precision, real-time monitoring, and expert-led interventions. In the ICU, a team of professionals continuously tracks a wide array of biomarkers, interprets their trends, and delivers timely care with orchestration and rigor. Y et, this reactive strength has come at the expense of a deeper, more continuous engagement with health as it unfolds in everyday life. This limitation was first articulated in the early calls for precision and P4 medicine, which envisioned predictive, personalized, preventive, and participatory models of care that would complement traditional clinical practice [1, 2, 3]. This imbalance is starkly captured in what we call the "8759 vs. 1" paradox: an individual spends 8759 hours each year outside the clinical setting, making decisions that shape their health -- while barely an hour is spent in direct consultation with care providers. During those other hours, health is continuously influenced by behavior, environment, emotion, and social context. Y et, our existing computing systems remain fixated on the one hour, neglecting the remaining 8759. 1


Cluster and Aggregate: Face Recognition with Large Probe Set Supplementary Material

Neural Information Processing Systems

The number of layers L in CN is equal to 2. For recent SoT A backbone models, the performance is saturated above 98 .5 . The performance gain is observed in both backbones. As the probe size increases, the role of a feature fusion model also increases. The relative performance gain for Fig.1 c) is calculated as We measured the FPS with Nvidia RTX3090. When a few samples' contribution is larger than the others Lower entropy value tells you that the cluster features are deviating from a simple average of all samples.


OpenAI Designed GPT-5 to Be Safer. It Still Outputs Gay Slurs

WIRED

OpenAI is trying to make its chatbot less annoying with the release of GPT-5. And I'm not talking about adjustments to its synthetic personality that many users have complained about. Before GPT-5, if the AI tool determined it couldn't answer your prompt because the request violated OpenAI's content guidelines, it would hit you with a curt, canned apology. Now, ChatGPT is adding more explanations. OpenAI's general model spec lays out what is and isn't allowed to be generated.


Mixture of Nested Experts: Adaptive Processing of Visual Tokens

Neural Information Processing Systems

The visual medium (images and videos) naturally contains a large amount of information redundancy, thereby providing a great opportunity for leveraging efficiency in processing. While Vision Transformer (ViT) based models scale effectively to large data regimes, they fail to capitalize on this inherent redundancy, leading to higher computational costs. Mixture of Experts (MoE) networks demonstrate scalability while maintaining same inference-time costs, but they come with a larger parameter footprint. We present Mixture of Nested Experts (MoNE), which utilizes a nested structure for experts, wherein individual experts fall on an increasing compute-accuracy curve. Given a compute budget, MoNE learns to dynamically choose tokens in a priority order, and thus redundant tokens are processed through cheaper nested experts.