Goto

Collaborating Authors

 credibility


Ex-ESPN host Cindy Brunson embarrasses herself trying to discredit Caitlin Clark

FOX News

Diamondbacks manager Torey Lovullo apologizes to Chicago Cubs after sign stealing controversy: 'bush league' Caitlin Clark helps Fever deliver WNBA's largest regular-season TV audience in 29 years Fans express concern over Tiger Woods' appearance in rare video since March DUI, rehab Commanders' Jayden Daniels deflects question on LSU cease-and-desist letter after social media trolling wave Female Little Leaguer is smashing records, CBS's new college football reporter & heartbreaking sports moments! Hot take alert: ESPN getting rid of'BottomLine' score ticker is a travesty, not a reason to celebrate Livvy Dunne selflessly passes along three keys to'Baywatch' slow motion run to NFL quarterback's girlfriend Shirin Yadegar: Let's be really clear about what Hasan Piker is messaging Iran doesn't have any significant attacks left that threaten anything: Ret Marine colonel Iran doesn't have any significant attacks left that threaten anything: Ret Marine colonel The radical left is no longer the'fringe,' Andrew Kolvet warns GOP needs'better messaging' on AI data centers: Tomi Lahren GOP needs'better messaging' on AI data centers: Tomi Lahren Meta acted on greed that cost kids' lives, says'Scrolling 2 Death' podcast host Meta acted on greed that cost kids' lives, says'Scrolling 2 Death' podcast host Brunson dismissed Clark's 37-point game due to the margin but praised Angel Reese's rebound record in a similar blowout WNBA fans outside Gainbridge Fieldhouse in Indiana talk to OutKick following the controversy stirred up by DiJonai Carrington over her White privilege remark. We previously discussed the all-time low the sports media reached in defending DiJonai Carrington after she clotheslined Sophie Cunningham and then posted White privilege from the locker room following her ejection. Brunson claimed Carrington did not deserve a Flagrant 2. Instead, she argued that Cunningham should have been ejected -- not the player who clotheslined her. Brunson also shared a post claiming white WNBA players benefit from reverse DEI .


Anyone can fake a scientific image with AI, tricking even academic journals – and undermining trust in science

AIHub

A photograph of Earth glowing in deep space, the Moon's cratered horizon stretching across its foreground, caught many people's eyes in April 2026. Astronauts captured the image while aboard NASA's Artemis II mission, and like the famous Apollo 8 "Earthrise" image, the picture felt instantly real and inspiring for many. But when almost anyone can fabricate a visually similar image in seconds from a text prompt using artificial intelligence, how do people decide which image is real? The proliferation of AI-generated science images in public spaces is not simply a misinformation problem. As a researcher who studies visual science communication and public trust, I believe it also contributes to a crisis of trust in science in the age of AI, and the tools scientists have long relied on to establish visual credibility are losing their grip.


The Download: metric weaknesses and AI elephant warnings

MIT Technology Review

Plus: The US has allowed Anthropic to release Mythos 5 to "trusted" orgs. There are plenty of useful things a metric can reveal. There are even more that it can obscure or corrupt. Like a lot of people bitten by the self-quantifying bug, I started gathering personal data to pursue a nebulous collection of goals and desires. I wanted to feel better physically and emotionally, get outside more, and bring order to the messiness and uncertainty of my daily existence. But external metrics and data can never capture what's truly important.


Position: Benchmarking is Broken - Don't Let AI be Its Own Judge

Neural Information Processing Systems

The meteoric rise of Artificial Intelligence (AI), with its rapidly expanding market capitalization, presents both transformative opportunities and critical challenges. Chief among these is the urgent need for a new, unified paradigm for trustworthy evaluation, as current benchmarks increasingly reveal critical vulnerabilities. Issues like data contamination and selective reporting by model developers fuel hype, while inadequate data quality control can lead to biased evaluations that, even if unintentionally, may favor specific approaches. As a flood of participants enters the AI space, this Wild West of assessment makes distinguishing genuine progress from exaggerated claims exceptionally difficult. Such ambiguity blurs scientific signals and erodes public confidence, much as unchecked claims would destabilize financial markets reliant on credible oversight from agencies like Moody's.In high-stakes human examinations (e.g., SAT, GRE), substantial effort is devoted to ensuring fairness and credibility; why settle for less in evaluating AI, especially given its profound societal impact? This position paper argues that a laissez-faire approach is untenable. For true and sustainable AI advancement, we call for a paradigm shift to a unified, live, and quality-controlled benchmarking framework--robust by construction rather than reliant on courtesy or goodwill.


The Download: Musk v. Altman week 3, and Trump's tech trading

MIT Technology Review

Musk v. Altman week 3: Musk and Altman traded blows over each other's credibility. Now the jury will pick a side. In the final week of the Musk v. Altman trial, lawyers attacked the credibility of the two tech leaders. Sam Altman was accused of lying and self-dealing, while Elon Musk was portrayed as a power-seeker trying to control artificial general intelligence. The case unearthed new details about the two arch-rivals and OpenAI's contested nonprofit status, as well as a golden trophy of a donkey's ass awarded to an employee who challenged Musk. Michelle Kim, who's also a lawyer, has been in court throughout the Musk v. Altman trial.


Reddit's human content wins amid the AI flood

BBC News

Reddit's human content wins amid the AI flood For Ines Tan there's one particular site she turns to again and again for advice - and that's Reddit. Tan, who works in communications, regularly jumps on the site for skincare advice, to view reactions to shows she watches, such as The Traitors, and for help planning her upcoming wedding in May. It's a very empathetic place, she says of Reddit. For my wedding, I've found help emotionally, logistically and inspiration-wise. Tan believes people are consulting the online discussion platform more as they're craving human interaction in the world of increasing AI slop.



Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines

arXiv.org Artificial Intelligence

LLM-based Search Engines (LLM-SEs) introduces a new paradigm for information seeking. Unlike Traditional Search Engines (TSEs) (e.g., Google), these systems summarize results, often providing limited citation transparency. The implications of this shift remain largely unexplored, yet raises key questions regarding trust and transparency. In this paper, we present a large-scale empirical study of LLM-SEs, analyzing 55,936 queries and the corresponding search results across six LLM-SEs and two TSEs. We confirm that LLM-SEs cites domain resources with greater diversity than TSEs. Indeed, 37% of domains are unique to LLM-SEs. However, certain risks still persist: LLM-SEs do not outperform TSEs in credibility, political neutrality and safety metrics. Finally, to understand the selection criteria of LLM-SEs, we perform a feature-based analysis to identify key factors influencing source choice. Our findings provide actionable insights for end users, website owners, and developers.


VP-AutoTest: A Virtual-Physical Fusion Autonomous Driving Testing Platform

arXiv.org Artificial Intelligence

The rapid development of autonomous vehicles has led to a surge in testing demand. Traditional testing methods, such as virtual simulation, closed-course, and public road testing, face several challenges, including unrealistic vehicle states, limited testing capabilities, and high costs. These issues have prompted increasing interest in virtual-physical fusion testing. However, despite its potential, virtual-physical fusion testing still faces challenges, such as limited element types, narrow testing scope, and fixed evaluation metrics. To address these challenges, we propose the Virtual-Physical Testing Platform for Autonomous Vehicles (VP-AutoTest), which integrates over ten types of virtual and physical elements, including vehicles, pedestrians, and roadside infrastructure, to replicate the diversity of real-world traffic participants. The platform also supports both single-vehicle interaction and multi-vehicle cooperation testing, employing adversarial testing and parallel deduction to accelerate fault detection and explore algorithmic limits, while OBU and Redis communication enable seamless vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) cooperation across all levels of cooperative automation. Furthermore, VP-AutoTest incorporates a multidimensional evaluation framework and AI-driven expert systems to conduct comprehensive performance assessment and defect diagnosis. Finally, by comparing virtual-physical fusion test results with real-world experiments, the platform performs credibility self-evaluation to ensure both the fidelity and efficiency of autonomous driving testing. Please refer to the website for the full testing functionalities on the autonomous driving public service platform OnSite:https://www.onsite.com.cn.


Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings

arXiv.org Artificial Intelligence

Recent advances enable Large Language Models (LLMs) to generate AI personas, yet their lack of deep contextual, cultural, and emotional understanding poses a significant limitation. This study quantitatively compared human responses with those of eight LLM-generated social personas (e.g., Male, Female, Muslim, Political Supporter) within a low-resource environment like Bangladesh, using culturally specific questions. Results show human responses significantly outperform all LLMs in answering questions, and across all matrices of persona perception, with particularly large gaps in empathy and credibility. Furthermore, LLM-generated content exhibited a systematic bias along the lines of the ``Pollyanna Principle'', scoring measurably higher in positive sentiment ($Φ_{avg} = 5.99$ for LLMs vs. $5.60$ for Humans). These findings suggest that LLM personas do not accurately reflect the authentic experience of real people in resource-scarce environments. It is essential to validate LLM personas against real-world human data to ensure their alignment and reliability before deploying them in social science research.