Goto

Collaborating Authors

 mastodon


Fedivertex: a Graph Dataset based on Decentralized Social Networks for Trustworthy Machine Learning

arXiv.org Artificial Intelligence

Decentralized machine learning - where each client keeps its own data locally and uses its own computational resources to collaboratively train a model by exchanging peer-to-peer messages - is increasingly popular, as it enables better scalability and control over the data. A major challenge in this setting is that learning dynamics depend on the topology of the communication graph, which motivates the use of real graph datasets for benchmarking decentralized algorithms. Unfortunately, existing graph datasets are largely limited to for-profit social networks crawled at a fixed point in time and often collected at the user scale, where links are heavily influenced by the platform and its recommendation algorithms. The Fediverse, which includes several free and open-source decentralized social media platforms such as Mastodon, Misskey, and Lemmy, offers an interesting real-world alternative. We introduce Fedivertex, a new dataset of 182 graphs, covering seven social networks from the Fediverse, crawled weekly over 14 weeks. We release the dataset along with a Python package to facilitate its use, and illustrate its utility on several tasks, including a new defederation task, which captures a process of link deletion observed on these networks.


Characterizing LLM-driven Social Network: The Chirper.ai Case

arXiv.org Artificial Intelligence

Large language models (LLMs) demonstrate the ability to simulate human decision-making processes, enabling their use as agents in modeling sophisticated social networks, both offline and online. Recent research has explored collective behavioral patterns and structural characteristics of LLM agents within simulated networks. However, empirical comparisons between LLM-driven and human-driven online social networks remain scarce, limiting our understanding of how LLM agents differ from human users. This paper presents a large-scale analysis of Chirper.ai, an X/Twitter-like social network entirely populated by LLM agents, comprising over 65,000 agents and 7.7 million AI-generated posts. For comparison, we collect a parallel dataset from Mastodon, a human-driven decentralized social network, with over 117,000 users and 16 million posts. We examine key differences between LLM agents and humans in posting behaviors, abusive content, and social network structures. Our findings provide critical insights into the evolving landscape of online social network analysis in the AI era, offering a comprehensive profile of LLM agents in social simulations.


A Simulation System Towards Solving Societal-Scale Manipulation

arXiv.org Artificial Intelligence

The rise of AI-driven manipulation poses significant risks to societal trust and democratic processes. Yet, studying these effects in real-world settings at scale is ethically and logistically impractical, highlighting a need for simulation tools that can model these dynamics in controlled settings to enable experimentation with possible defenses. We present a simulation environment designed to address this. We elaborate upon the Concordia framework that simulates offline, 'real life' activity by adding online interactions to the simulation through social media with the integration of a Mastodon server. We improve simulation efficiency and information flow, and add a set of measurement tools, particularly longitudinal surveys. We demonstrate the simulator with a tailored example in which we track agents' political positions and show how partisan manipulation of agents can affect election results.


Safeguarding Decentralized Social Media: LLM Agents for Automating Community Rule Compliance

arXiv.org Artificial Intelligence

Ensuring content compliance with community guidelines is crucial for maintaining healthy online social environments. However, traditional human-based compliance checking struggles with scaling due to the increasing volume of user-generated content and a limited number of moderators. Recent advancements in Natural Language Understanding demonstrated by Large Language Models unlock new opportunities for automated content compliance verification. This work evaluates six AI-agents built on Open-LLMs for automated rule compliance checking in Decentralized Social Networks, a challenging environment due to heterogeneous community scopes and rules. Analyzing over 50,000 posts from hundreds of Mastodon servers, we find that AI-agents effectively detect non-compliant content, grasp linguistic subtleties, and adapt to diverse community contexts. Most agents also show high inter-rater reliability and consistency in score justification and suggestions for compliance. Human-based evaluation with domain experts confirmed the agents' reliability and usefulness, rendering them promising tools for semi-automated or human-in-the-loop content moderation systems.


An Innovative Tool for Uploading/Scraping Large Image Datasets on Social Networks

arXiv.org Artificial Intelligence

Nowadays, people can retrieve and share digital information in an increasingly easy and fast fashion through the well-known digital platforms, including sensitive data, inappropriate or illegal content, and, in general, information that might serve as probative evidence in court. Consequently, to assess forensics issues, we need to figure out how to trace back to the posting chain of a digital evidence (e.g., a picture, an audio) throughout the involved platforms -- this is what Digital (also Forensics) Ballistics basically deals with. With the entry of Machine Learning as a tool of the trade in many research areas, the need for vast amounts of data has been dramatically increasing over the last few years. However, collecting or simply find the "right" datasets that properly enables data-driven research studies can turn out to be not trivial in some cases, if not extremely challenging, especially when it comes with highly specialized tasks, such as creating datasets analyzed to detect the source media platform of a given digital media. In this paper we propose an automated approach by means of a digital tool that we created on purpose. The tool is capable of automatically uploading an entire image dataset to the desired digital platform and then downloading all the uploaded pictures, thus shortening the overall time required to output the final dataset to be analyzed.


'World of Warcraft' Has a Lot to Teach the Twitter Clones

WIRED

Another week, another catastrophic failure of policy at Twitter that's being eagerly exploited by its myriad competitors--and they truly are myriad. And yet, in spite of the momentary success of some of these platforms--Threads has gotten over 70 million signups as of this writing--none has quite ascended to the lofty heights of Twitter's influence at its height, where it seemed, for good or for ill (let's be honest, mostly ill), to be at the heart of every conversation among our world's epistemic elites. To understand why, we have to go to Azeroth. Black Mirror creator Charlie Brooker once, tongue-in-cheek, called Twitter the best video game of all time, likening it to the then-still-popular wave of Massively Multiplayer Online Roleplaying Games, or MMORPGs, that were led by titles like World of Warcraft. Aside from the obvious connections--adopting an online persona in a gamified system entirely governed by earned metrics--we can also look at the fact that Twitter, like World of Warcraft, is surrounded by failed imitators.


Erase browser history: can AI reset the browser battle? - The Verge

#artificialintelligence

I think that's a really interesting question. What is the nature of the community around Mastodon, right? When we think about it, how much is the protocol itself, and how much is actually the community of people engaging with it, building things, and trying to do something new? The protocol itself is a distributed protocol, and they take time and energy and stuff to build. But the real success also needs a set of people who are interested enough to do something different.


Smashing Security podcast #307: ChatGPT and the Minister for Foreign Affairs • Graham Cluley

#artificialintelligence

Could a senior Latvian politician really be responsible for scamming hundreds of "mothers-of-two" in the UK? (Probably not, despite Graham's theories…) And should we be getting worried about the AI wonder that is ChatGPT? All this and more is discussed in the latest edition of the "Smashing Security" podcast by computer security veterans Graham Cluley and Carole Theriault. You can help the podcast by telling your friends and colleagues about "Smashing Security", and leaving us a review on Apple Podcasts or Podchaser. Become a supporter via Patreon or Apple Podcasts for ad-free episodes and our early-release feed! Follow the show on Twitter at @SmashinSecurity, or on Mastodon, on the Smashing Security subreddit, or visit our website for more episodes.


Cybersecurity Winter is coming

#artificialintelligence

The scip Blog Digest is a monthly released summary of the most important, thrilling and crazy posts from the international blogosphere. While reading this digest it remains very easy to keep up to date with the events of cybersecurity and advanced technology. Follow our team on Twitter and the company on LinkedIn, to get the most recent news. Base editing: Revolutionary therapy clears girl's incurable cancer Cyber attacks set to become'uninsurable', says Zurich chief Data breach at Social Blade confirmed. Do Users Write More Insecure Code with AI Assistants?


What the midterm madness means for startups

#artificialintelligence

Welcome to Startups Weekly, a nuanced take on this week's startup news and trends. To get this in your inbox, subscribe here. It's Kyle, filling in this issue for Natasha, who's taking a much needed break from the news cycle (and the spectacle that's become Twitter). While it's my first Startups Weekly column, you've likely seen me on TC here and there, covering chiefly venture, AI and enterprise-related items. It's a real pleasure to round up this week's startup news -- partially because it doesn't center around Musk shenanigans.