Goto

Collaborating Authors

 specialism


Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization James Oldfield

Neural Information Processing Systems

An important corollary of successful task decomposition amongst experts is that layers are easier to debug and edit. Biased or unsafe behaviors can be better localized to specific experts' subcomputation, facilitating manual correction or surgery in a way that minimally affects the other functionality of the network. Addressing such behaviors is particularly crucial in the context of foundation models; being often fine-tuned as black boxes pre-trained on unknown, potentially imbalanced data distributions.



A Practical Guide to Interpretable Role-Based Clustering in Multi-Layer Financial Networks

arXiv.org Artificial Intelligence

Understanding the functional roles of financial institutions within interconnected markets is critical for effective supervision, systemic risk assessment, and resolution planning. We propose an interpretable role-based clustering approach for multi-layer financial networks, designed to identify the functional positions of institutions across different market segments. Our method follows a general clustering framework defined by proximity measures, cluster evaluation criteria, and algorithm selection. We construct explainable node embeddings based on egonet features that capture both direct and indirect trading relationships within and across market layers. Using transaction-level data from the ECB's Money Market Statistical Reporting (MMSR), we demonstrate how the approach uncovers heterogeneous institutional roles such as market intermediaries, cross-segment connectors, and peripheral lenders or borrowers. The results highlight the flexibility and practical value of role-based clustering in analyzing financial networks and understanding institutional behavior in complex market structures.


Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization

arXiv.org Artificial Intelligence

The Mixture of Experts (MoE) paradigm provides a powerful way to decompose dense layers into smaller, modular computations often more amenable to human interpretation, debugging, and editability. However, a major challenge lies in the computational cost of scaling the number of experts high enough to achieve fine-grained specialization. In this paper, we propose the Multilinear Mixture of Experts ($\mu$MoE) layer to address this, focusing on vision models. $\mu$MoE layers enable scalable expert specialization by performing an implicit computation on prohibitively large weight tensors entirely in factorized form. Consequently, $\mu$MoEs (1) avoid the restrictively high inference-time costs of 'soft' MoEs, yet (2) do not inherit the training issues of the popular 'sparse' MoEs' discrete (non-differentiable) expert routing. We present both qualitative and quantitative evidence that scaling $\mu$MoE layers when fine-tuning foundation models for vision tasks leads to more specialized experts at the class-level, further enabling manual bias correction in CelebA attribute classification. Finally, we show qualitative results demonstrating the expert specialism achieved when pre-training large GPT2 and MLP-Mixer models with parameter-matched $\mu$MoE blocks at every layer, maintaining comparable accuracy. Our code is available at: https://github.com/james-oldfield/muMoE.


Will 2021 Be The Year That AI Finally Scales?

#artificialintelligence

The gap between the promise of Artificial Intelligence (AI) and its implementation in practice has never been greater than it was in 2020. There were clearly some major milestone AI achievements last year. Take Google DeepMind's AlphaFold, which was shown to accurately predict 3D models of protein structures, paving the way for groundbreaking research across every field of biology. Or in June, when a beta version of GTP3 was publicly released by Microsoft - an incredibly sophisticated model capable of almost any language task, including writing in the style of Chaucer, and even basic coding. Yet outside of these tech giants, AI adoption remains in exploratory stages for the vast majority of enterprises and a long way off becoming an integral part of day to day business. At this point in the AI adoption cycle, many enterprises hold an untenable position as long as they fail to appreciate the enormous potential for embedding machine learning into their products and business processes.


Here's what you need to know before going under the robo-knife

Daily Mail - Science & tech

First, you are strapped from the chest upwards on to the table, with your feet hoisted into stirrups. The table is swung down backwards, so you are tilted, head-down, at an angle of 45 degrees. Then a machine, known by some surgeons as'the 800lb gorilla', can get to work. It sounds so medieval, but this is the most modern of surgical techniques -- robotic surgery. The extraordinary posture, known as the steep Trendelenburg, is necessary to position the patient precisely so the robot arms can reach inside them.


Let's stop kidding ourselves about SEO and artificial intelligence

#artificialintelligence

Now more than ever the SEO industry looks to be mirroring that of politics, with sound bites and buzzwords the order of the day. Grand lanyard-heavy conferences, emphatic speeches thrown across eager audiences, the sharing of likeminded ideals, and the spinning of content -- it's all there. I'm not even talking about SEO: The Movie, where link building meets Wolf of Wall Street meets the losing players of The Social Network. Over the past few months I've seen a large amount of respected agencies talk about offering AIO specialisms to their clients. But there's a snag about offering AIO to your clients, because AIO doesn't exist.