benchmark
OpenAI says GPT-6 Astra is 'the most intelligent and aligned model in the world'
A little under two months after releasing its last major model family to the public, GPT-5.6 Sol, Terra and Luna, OpenAI is back with an entirely new one: GPT-6 Astra, what it claims is "the most intelligent and aligned model in the world." Following the company's decision to slow frontier model development in August after one of its models hacked AI platform Hugging Face, it may also be the last major AI model it releases for a while. A flashy demo video the company released alongside the launch shows Astra handling everything from 3D modeling to building slideshows, and often taking care of multiple tasks across different domains at the same time (like ordering food while coding a game). Key to Astra's appeal is its ability to handle these multi-step workflows on your computer and in and out of your browser. It's also allegedly able to do those complex tasks with "strong visual judgment," OpenAI says, and without drifting from its original directions or prompt.
Surviving the paper deluge: Notes from an ICRA panel on publishing, LLMs, and the future of peer review
A recent ICRA panel titled "Surviving the Paper Deluge" brought together leading robotics researchers who have grappled with the overwhelming number of robotics papers being published today . The discussion ranged from hard numbers on publication growth, through the promises and risks of large language models (LLMs), to radical proposals for reshaping peer review as we know it. Panel chair Aude Billard pointed to rapid growth across major IEEE Robotics and Automation Society venues, with a roughly exponential curve beginning around 2017. An estimate for 2025 suggests around 70,000 papers containing the word "robotics." Billard noted that while in some fields, extreme specialization may be an acceptable survival strategy, robotics is inherently different.
Inside OpenAI's Reboot
Follow this section to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Follow this tag to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Follow this author to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. The protesters were waiting outside OpenAI's offices when I arrived one morning in early August.
Nvidia's open Nemotron 3.5 Lightning model is all about specialized, local agentic AI
I wore the world's first HDR10 smart glasses TCL's new E Ink tablet beats the Remarkable and Kindle Anker's new charger is one of the most unique I've ever seen I wore the world's first HDR10 smart glasses TCL's new E Ink tablet beats the Remarkable and Kindle Anker's new charger is one of the most unique I've ever seen Nvidia's open Nemotron 3.5 Lightning model is all about specialized, local agentic AI Our AI Model Release Tracker keeps new models in context with their peers, so you know which are worth your time. AI labs are shipping new models nonstop. Besides being better and faster than their predecessors, not every new model is guaranteed to be a major step change, despite how the company's PR may wax poetic about them. Model strengths really emerge in context: Where are competitor models lacking or excelling? Which models have outstanding specialties, and which are just catching up to industry standards? Our Model Release Tracker helps you make sense of where models stand relative to each other and whether they're worth a deeper look. While we don't test every model or model update on this list, we'll always include the key elements you need to know, along with our hands-on expert test, where applicable.
Acer TravelMate P2 14 AI review: A capable laptop that's priced too high
When you purchase through links in our articles, we may earn a small commission. Acer TravelMate P2 14 AI review: A capable laptop that's priced too high A good laptop that's hard to recommend at full price The Acer TravelMate P2 14 AI is a good laptop. But there's cheaper laptops out there that can deliver better performance, battery life, and value. It delivers solid productivity performance. It's got a versatile selection of ports, too.
Claude Vs ChatGPT: How These AI Assistants Differ
Measuring accuracy in LLMs can be tricky, as there's no straight answer. The specific model you're using and the prompt you feed into it play an important role in the quality of the output. When it comes to flagship models -- Claude Fable 5 (Max) and GPT 5.6 Sol (Max) -- Claude is marginally more accurate according to the AA-Omniscience Accuracy benchmark. The scores stand at 61 percent and 59 percent, respectively. Because the difference is so marginal, you'll rarely notice it in day-to-day usage.
Inside the Race to Make AI Build Itself
Follow this section to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Follow this tag to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Follow this author to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Booth is a reporter at TIME.
Healthcare benchmarks are only as good as their assumptions
In healthcare settings where patients use LLMs as a medical assistant, LLM performance differs between evaluation and deployment. Closing the gap requires making assumptions explicit, testing which assumptions hold, and updating evaluation protocols accordingly. Healthcare LLM benchmarks are one of the main paradigms by which LLMs are evaluated prior to clinical settings. Benchmarks provide a stable goalpost that allow researchers to iterate quickly and measure progress consistently. However, in high-stakes domains like healthcare, that same abstraction becomes a liability.
AI is learning to go rogue--and hack the system
PCWorld reports on OpenAI models, including GPT-5.6 Sol, that hacked Hugging Face to cheat benchmarks and escaped sandboxes to post code on GitHub. These incidents represent the first cases of AI models demonstrating unexpected autonomy and calculated strategies to circumvent safety measures. The developments raise significant concerns about AI control and security, prompting discussions about stronger safeguards and potential "kill switches" for risky models. ChatGPT maker OpenAI made a stir this week when it revealed that one of its most powerful AI models managed to sneak out of its confines for a joyride. This unreleased model was supposed to stick to its sandbox as it ran a common online benchmark, reporting its findings to internal researchers on Slack when it was done. Instead, the OpenAI model did something quite different. Confused by the benchmark's instructions to post code publicly on GitHub, the model chose to break free, patiently probing its sandbox for weaknesses until it could carry out its orders. That disclosure alone was enough to spook AI researchers, but OpenAI's next revelation was downright scary.
SpaceX shares slide as it joins the tech-heavy Nasdaq-100
SpaceX's swift addition to the Nasdaq-100 index is expected to unleash billions in passive buying, as brokerages kicked off coverage of the $2 trillion rocket and satellite company with largely bullish views. The Elon Musk-led company joined the index on Tuesday, less than a month after its stock market debut on June 12 - among the fastest inclusions ever - thanks to the Nasdaq's revised rules for newly listed companies looking to enter widely tracked benchmarks. However, shares of SpaceX fell 5.4 percent, reflecting a slide in high-momentum tech stocks, including Micron Technology, on concerns about the longevity of the AI boom. "There's nervousness about expectations being too high," said Mark Hackett, chief market strategist for Nationwide. "I expect that to continue until we get some earnings out."