Government
L-CiteEval: Do Long-Context Models Truly Leverage Context for Responding?
Tang, Zecheng, Zhou, Keyan, Li, Juntao, Ji, Baibei, Hou, Jianye, Zhang, Min
Long-context models (LCMs) have made remarkable strides in recent years, offering users great convenience for handling tasks that involve long context, such as document summarization. As the community increasingly prioritizes the faithfulness of generated results, merely ensuring the accuracy of LCM outputs is insufficient, as it is quite challenging for humans to verify the results from the extremely lengthy context. Yet, although some efforts have been made to assess whether LCMs respond truly based on the context, these works either are limited to specific tasks or heavily rely on external evaluation resources like GPT4.In this work, we introduce L-CiteEval, a comprehensive multi-task benchmark for long-context understanding with citations, aiming to evaluate both the understanding capability and faithfulness of LCMs. L-CiteEval covers 11 tasks from diverse domains, spanning context lengths from 8K to 48K, and provides a fully automated evaluation suite. Through testing with 11 cutting-edge closed-source and open-source LCMs, we find that although these models show minor differences in their generated results, open-source models substantially trail behind their closed-source counterparts in terms of citation accuracy and recall. This suggests that current open-source LCMs are prone to responding based on their inherent knowledge rather than the given context, posing a significant risk to the user experience in practical applications. We also evaluate the RAG approach and observe that RAG can significantly improve the faithfulness of LCMs, albeit with a slight decrease in the generation quality. Furthermore, we discover a correlation between the attention mechanisms of LCMs and the citation generation process.
A large-scale operational study of fingerprint quality and demographics
Galbally, Javier, Cepilovs, Aleksandrs, Blanco-Gonzalo, Ramon, Ormiston, Gillian, Miguel-Hurtado, Oscar, Racz, Istvan Sz.
Abstract--Even though a few initial works have shown on small sets of data some level of bias in the performance of fingerprint recognition technology with respect to certain demographic groups, there is still not sufficient evidence to understand the impact that certain factors such as gender, age or finger-type may have on fingerprint quality and, in turn, also on fingerprint matching accuracy. The present work addresses this still under researched topic, on a large-scale database of operational data containing 10-print impressions of almost 16,000 subjects. The results reached provide further insight into the dependency of fingerprint quality and demographics, and show that there in fact exists a certain degree of performance variability in fingerprint-based recognition systems for different segments of the population. Based on the experimental evaluation, the work points out new observations based on data-driven evidence, provides plausible hypotheses to explain such observations, and concludes with potential follow-up actions that can help to reduce the observed fingerprint quality differences. This way, the current paper can be considered as a contribution to further increase the algorithmic fairness and equality of biometric technology. "It's not the size of the dog in the fight, it's the size of demographic group, why do some segments of the population the fight in the dog" - Mark Twain However, with the exception of a few studies, comprise more information than those of young children or this inconsistency in the recognition rates has been mainly elders? Why do each of the fingers (including the thumb) of observed on small-to-medium databases under laboratory the hand provide different accuracy performance in fingerprint conditions and, therefore, it is difficult to quantify to what recognition systems?
Multiscale Semi-Markov Dynamics for Intracortical Brain-Computer Interfaces
Daniel Milstein, Jason Pacheco, Leigh Hochberg, John D. Simeral, Beata Jarosiewicz, Erik Sudderth
Intracortical brain-computer interfaces (iBCIs) have allowed people with tetraplegia to control a computer cursor by imagining the movement of their paralyzed arm or hand. State-of-the-art decoders deployed in human iBCIs are derived from a Kalman filter that assumes Markov dynamics on the angle of intended movement, and a unimodal dependence on intended angle for each channel of neural activity. Due to errors made in the decoding of noisy neural data, as a user attempts to move the cursor to a goal, the angle between cursor and goal positions may change rapidly. We propose a dynamic Bayesian network that includes the on-screen goal position as part of its latent state, and thus allows the person's intended angle of movement to be aggregated over a much longer history of neural activity. This multiscale model explicitly captures the relationship between instantaneous angles of motion and long-term goals, and incorporates semi-Markov dynamics for motion trajectories. We also introduce a multimodal likelihood model for recordings of neural populations which can be rapidly calibrated for clinical applications. In offline experiments with recorded neural data, we demonstrate significantly improved prediction of motion directions compared to the Kalman filter. We derive an efficient online inference algorithm, enabling a clinical trial participant with tetraplegia to control a computer cursor with neural activity in real time. The observed kinematics of cursor movement are objectively straighter and smoother than prior iBCI decoding models without loss of responsiveness.
Buttigieg's message on restricting civilian drones near Hurricane Helene damage prompts outcry, clarification
Fox News contributor Marc Thiessen joined'Fox & Friends' to discuss why President Biden and Vice President Harris are facing scrutiny for their response to Hurricane Helene. The U.S. Department of Transportation (DOT) clarified a message that warned civilian drone pilots not to fly near Hurricane Helene recovery and rescue efforts -- or risk penalty, fines or "criminal prosecution" -- after facing intense backlash online. Reached by Fox News Digital, a DOT spokesperson said civilian drone pilots are permitted and are assisting in rescue and recovery efforts, and previous "temporary flight restrictions" have since been lifted. Some X users -- collectively with millions of followers -- reacted adversely to a message addressed to drone pilots and with accompanying video from Transportation Secretary Pete Buttigieg shared by the department earlier this week. The message and video argued the restrictions would prohibit civilian volunteers from legally searching for victims or survivors when response time matters most or capturing their own footage of the disaster.
Judge blocks California law that targeted deepfake campaign ads
With deepfake video and audio making their way into political campaigns, California enacted its toughest restrictions yet in September: a law prohibiting political ads within 120 days of an election that include deceptive, digitally generated or altered content unless the ads are labeled as "manipulated." On Wednesday, a federal judge temporarily blocked the law, saying it violated the 1st Amendment. Other laws against deceptive campaign ads remain on the books in California, including one that requires candidates and political action committees to disclose when ads are using artificial intelligence to create or substantially alter content. But the preliminary injunction granted against Assembly Bill 2839 means that there will be no broad prohibition against individuals using artificial intelligence to clone a candidate's image or voice and portraying them falsely without revealing that the images or words are fake. The injunction was sought by Christopher Kohls, a conservative commentator who has created a number of deepfake videos satirizing Democrats, including the party's presidential nominee, Vice President Kamala Harris.
An Empirical Study on The Properties of Random Bases for Kernel Methods
Kernel machines as well as neural networks possess universal function approximation properties. Nevertheless in practice their ways of choosing the appropriate function class differ. Specifically neural networks learn a representation by adapting their basis functions to the data and the task at hand, while kernel methods typically use a basis that is not adapted during training. In this work, we contrast random features of approximated kernel machines with learned features of neural networks. Our analysis reveals how these random and adaptive basis functions affect the quality of learning. Furthermore, we present basis adaptation schemes that allow for a more compact representation, while retaining the generalization properties of kernel machines.
US Army is testing 'Lone Wolf' robot dog with AI-powered rifle in the Middle East
The US Army is closer to unleashing robots on the battlefield after sending one dubbed'Lone Wolf' to the Middle East. The robot dog features an AR-15/M16-pattern rifle on its back that is attached to an AI-powered rotating mount capable of spotting aerial targets. The armed machine was sent overseas for rehearsal drills at the Red Sands Integrated Experimentation Center in Saudi Arabia. The military shared a photo of Lone Wolf last week, showing a Korean-made Ghost Robotics Vision 60 Quadrupedal-Unmanned Ground Vehicle (Q-UGV) at an undisclosed location. The US Army recently carried out testing of a new war machine in the Middle East.
'Ghost Ship of the Pacific' rediscovered with underwater drones
An autonomous drone fleet overseen by Ocean Infinity has rediscovered the USS Stewart, the only US Navy destroyer ever captured by Japanese forces during World War II. The marine robotics company's trio of orange, 20-foot-long underwater robots found the historic vessel while mapping what is now the 1,286-square-mile Cordell Bank national marine sanctuary off the California coast. Also known as the "Ghost Ship of the Pacific," the 314-foot-long ship has spent the past 78 years resting roughly 3,500 feet below the ocean's surface, and appears to remain almost completely intact and upright. "This level of preservation is exceptional for a vessel of its age and makes it potentially one of the best-preserved examples of a US Navy'four-piper' destroyer known to exist," Maria Brown, superintendent for both Cordell Bank and Greater Farallones national marine sanctuaries, said in a statement to The New York Times on October 1. The USS Stewart's story is unique in US maritime history, making it one of the most sought-after wrecks for decades.
Concrete Dropout
Yarin Gal, Jiri Hron, Alex Kendall
Dropout is used as a practical tool to obtain uncertainty estimates in large vision models and reinforcement learning (RL) tasks. But to obtain well-calibrated uncertainty estimates, a grid-search over the dropout probabilities is necessary-- a prohibitive operation with large models, and an impossible one with RL. We propose a new dropout variant which gives improved performance and better calibrated uncertainties. Relying on recent developments in Bayesian deep learning, we use a continuous relaxation of dropout's discrete masks. Together with a principled optimisation objective, this allows for automatic tuning of the dropout probability in large models, and as a result faster experimentation cycles. In RL this allows the agent to adapt its uncertainty dynamically as more data is observed. We analyse the proposed variant extensively on a range of tasks, and give insights into common practice in the field where larger dropout probabilities are often used in deeper model layers.