Law
An Anthropologist LLM to Elicit Users' Moral Preferences through Role-Play
De Ninno, Gianluca, Inverardi, Paola, Belotti, Francesca
GPT can predict users' future decisions by analyzing narrative tables, with accuracy further improved when guided by an anthropological framework. Moreover, by integrating contextual knowledge and an interpretative lens into LLMs, this approach enhances AI explainability while ensuring a human-centric perspective in requirement elicitation. By asking GPT to generate a user profile, it becomes possible to directly assess what the model has understood about the user and how it represents them. Furthermore, since the model is not only tasked with predicting users' responses in new scenarios but also with justifying its choices, it is possible, on one hand, to understand the rationale behind the model's output and, on the other, to identify potential misalignments between the model's prediction and the user's actual values and preferences. This enables targeted interventions to improve alignment between the LLM and the user profile, creating a continuous feedback loop that involves both the user and the LLM trained to interpret data through an anthropological lens. The process strengthens the model's interpretability, ethical alignment, and predictive adaptability, thereby making AI systems more transparent and attuned to real-world human values. Ultimately, the approach lays the groundwork for AI assistants capable of recognizing and adapting to individuals' soft ethics and ethical decision-making process. B. Threat to V alidity We discuss threats to validity following the qualitative research framework proposed in [72]--namely, credibility, transferability, dependability, and confirmability.
Rethinking Reward Models for Multi-Domain Test-Time Scaling
Lee, Dong Bok, Lee, Seanie, Park, Sangwoo, Kang, Minki, Baek, Jinheon, Kim, Dongki, Wagner, Dominik, Jin, Jiongdao, Lee, Heejun, Bocklet, Tobias, Wang, Jinyu, Fu, Jingjing, Hwang, Sung Ju, Bian, Jiang, Song, Lei
The reliability of large language models (LLMs) during test-time scaling is often assessed with \emph{external verifiers} or \emph{reward models} that distinguish correct reasoning from flawed logic. Prior work generally assumes that process reward models (PRMs), which score every intermediate reasoning step, outperform outcome reward models (ORMs) that assess only the final answer. This view is based mainly on evidence from narrow, math-adjacent domains. We present the first unified evaluation of four reward model variants, discriminative ORM and PRM (\DisORM, \DisPRM) and generative ORM and PRM (\GenORM, \GenPRM), across 14 diverse domains. Contrary to conventional wisdom, we find that (i) \DisORM performs on par with \DisPRM, (ii) \GenPRM is not competitive, and (iii) overall, \GenORM is the most robust, yielding significant and consistent gains across every tested domain. We attribute this to PRM-style stepwise scoring, which inherits label noise from LLM auto-labeling and has difficulty evaluating long reasoning trajectories, including those involving self-correcting reasoning. Our theoretical analysis shows that step-wise aggregation compounds errors as reasoning length grows, and our empirical observations confirm this effect. These findings challenge the prevailing assumption that fine-grained supervision is always better and support generative outcome verification for multi-domain deployment. We publicly release our code, datasets, and checkpoints at \href{https://github.com/db-Lee/Multi-RM}{\underline{\small\texttt{https://github.com/db-Lee/Multi-RM}}} to facilitate future research in multi-domain settings.
Generating Difficult-to-Translate Texts
Zouhar, Vilém, Xu, Wenda, Riley, Parker, Juraska, Juraj, Finkelstein, Mara, Freitag, Markus, Deutsch, Daniel
Machine translation benchmarks sourced from the real world are quickly obsoleted, due to most examples being easy for state-of-the-art translation models. This limits the benchmark's ability to distinguish which model is better or to reveal models' weaknesses. Current methods for creating difficult test cases, such as subsampling or from-scratch synthesis, either fall short of identifying difficult examples or suffer from a lack of diversity and naturalness. Inspired by the iterative process of human experts probing for model failures, we propose MT-breaker, a method where a large language model iteratively refines a source text to increase its translation difficulty. The LLM iteratively queries a target machine translation model to guide its generation of difficult examples. Our approach generates examples that are more challenging for the target MT model while preserving the diversity of natural texts. While the examples are tailored to a particular machine translation model during the generation, the difficulty also transfers to other models and languages.
Landcover classification and change detection using remote sensing and machine learning: a case study of Western Fiji
Gurjar, Yadvendra, Wan, Ruoni, Farahbakhsh, Ehsan, Chandra, Rohitash
As a developing country, Fiji is facing rapid urbanisation, which is visible in the massive development projects that include housing, roads, and civil works. In this study, we present machine learning and remote sensing frameworks to compare land use and land cover change from 2013 to 2024 in Nadi, Fiji. The ultimate goal of this study is to provide technical support in land cover/land use modelling and change detection. We used Landsat-8 satellite image for the study region and created our training dataset with labels for supervised machine learning. We used Google Earth Engine and unsupervised machine learning via k-means clustering to generate the land cover map. We used convolutional neural networks to classify the selected regions' land cover types. We present a visualisation of change detection, highlighting urban area changes over time to monitor changes in the map.
NASA goes dark hours before first look at interstellar object moving closer to Earth
Anguished Diddy clutches his head in his hands in first image from disastrous sentencing that's gone from bad to worse Mystery deepens over Hulk Hogan's death as his widow faces fresh anguish I'm no longer sleeping with my husband - and never will again, says MOLLY RYDDELL. I love him, but counted down the moments until he climaxed. Then I couldn't bear it any more and the truth spilled out... so many women feel the same Map shows where new strain of Covid is exploding in 19 states as sufferers are hit with'razor-blade' symptoms US military poised to seize ports and airfields in Venezuela as Trump strikes a fourth'narco-terrorist' boat Body count from Houston's bayous rises as serial killer whispers grip city and residents are told: 'Be vigilant' Selena Gomez's $1.3B fortune could create risks in Benny Blanco marriage despite his $50M success, experts reveal His daughter was warped into an ultra-woke monster and set fire to his life. Now, GOP state senator Jay Block fights back... and reveals the dark secrets she was desperate to hide Realtor with expensive ex-wife arrested over shocking $11.6m claims about how he was funding Palm Beach lifestyle The'middle-class kinks' saving marriages: Wives reveal the eight buzzy sex trends that revived their lagging libidos - including the fantasy husbands are secretly obsessed with Scientists discover key part of the brain that degrades in Alzheimer's... paving way for breakthrough therapies Lori Loughlin's estranged husband Mossimo Giannulli seen with mystery brunette amid shock split Manchester synagogue terrorist was on bail for alleged rape at the time of his rampage and was'struggling with debt' after'splitting up with his wife and young son' Scientists behind study linking Tylenol to autism accuse Trump of'spreading misinformation' Teresa Giudice thought she was going to'die' during panic attack on Special Forces... after she was called'stupid' NASA has gone dark just hours before humans get the closest look at the mysterious object barreling through our solar system . The interstellar object dubbed 3I/ATLAS will come within 18 million miles of Mars on October 3, its closest flyby of any planet this year.
Supplementary Material: CARLANE: A Lane Detection Benchmark for Unsupervised Domain Adaptation from Simulation to multiple Real-World Domains
Does the dataset contain all possible instances or is it a sample (not necessarily random) of instances from a larger set? If the dataset is a sample, then what is the larger set? Is the sample representative of the larger set (e.g., geographic coverage)? If so, please describe how this representativeness was validated/verified. If it is not representative of the larger set, please describe why not (e.g., to cover a more diverse range of instances,
Tesla sales jump as buyers scramble before EV tax credit expires
Tesla sales have surged in the third quarter as buyers in the United States rushed to take advantage of electric vehicle (EV) tax credits that were eliminated under President Donald Trump's sweeping tax bill passed this year. On Thursday, the automaker reported a 7.4 percent increase in sales compared with the same period last year as demand was driven by customers looking to buy before the credits officially expired at the end of September. Tesla also delivered 481,166 units of its Model 3 compact sedan and Model Y crossover in the quarter, well above Wall Street expectations. The Elon Musk-led carmaker frequently talked up the expiry of the tax credits, using it alongside discounts and financing deals to spur sales and leases of its EVs. Investors are worried because sales are now expected to slump as the $7,500 federal tax credit disappears.
Thousands of Americans can't get new jobs as government shutdown directly hits citizens
Overall, federal records revealed that over 2.7 million background checks are conducted every month for various reasons, including job applications. This also means air travel will likely suffer this month because the FAA Pilot Record Database is delayed and new pilots cannot start flights. FBI criminal checks, as well as civil court record searches, are needed for nurses, aides, and therapists, to slow hospital onboarding. 'Shutdowns cause delays in court record updates and verification gaps, slowing credentialing and risking non-compliance with HIPAA or state mandates,' the cyber expert explained. Trucking and shipping companies will also feel the impact of the shutdown because commercial drivers must undergo mandated checks by the Drug and Alcohol Clearinghouse and Pre-Employment Screening Program (PSP), which depend on federal IT updates and audits.