instituto
Instituto de Telecomunicações at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning
Attanasio, Giuseppe, Sannigrahi, Sonal, Peters, Ben, Martins, André F. T.
This paper presents the IT-IST submission to the IWSLT 2025 Shared Task on Instruction Following Speech Processing. We submit results for the Short Track, i.e., speech recognition, translation, and spoken question answering. Our model is a unified speech-to-text model that integrates a pre-trained continuous speech encoder and text decoder through a first phase of modality alignment and a second phase of instruction fine-tuning. Crucially, we focus on using small-scale language model backbones (< 2B) and restrict to high-quality, CC-BY data along with synthetic data generation to supplement existing resources.
MEL: Legal Spanish Language Model
Sánchez, David Betancur, García, Nuria Aldama, Jiménez, Álvaro Barbero, Nieto, Marta Guerrero, Morales, Patricia Marsà, Salas, Nicolás Serrano, Hernán, Carlos García, Coll, Pablo Haya, Ponsoda, Elena Montiel, Ibáñez, Pablo Calleja
Legal texts, characterized by complex and specialized terminology, present a significant challenge for Language Models. Adding an underrepresented language, such as Spanish, to the mix makes it even more challenging. While pre-trained models like XLM-RoBERTa have shown capabilities in handling multilingual corpora, their performance on domain specific documents remains underexplored. This paper presents the development and evaluation of MEL, a legal language model based on XLM-RoBERTa-large, fine-tuned on legal documents such as BOE (Bolet\'in Oficial del Estado, the Spanish oficial report of laws) and congress texts. We detail the data collection, processing, training, and evaluation processes. Evaluation benchmarks show a significant improvement over baseline models in understanding the legal Spanish language. We also present case studies demonstrating the model's application to new legal texts, highlighting its potential to perform top results over different NLP tasks.
3CEL: A corpus of legal Spanish contract clauses
García, Nuria Aldama, Morales, Patricia Marsà, Sánchez, David Betancur, Jiménez, Álvaro Barbero, Nieto, Marta Guerrero, Coll, Pablo Haya, Chozas, Patricia Martín, Ponsoda, Elena Montiel
Information extraction (IE) is defined as the NLP task that deals with the identification of particular pieces of information in unstructured documents [1, 2, 3]. In other words, the main objective of IE is to spot predefined relevant information in raw text. IE includes different subtypes depending on the nature of the information to be extracted. Thus, Named Entity Recognition (NER), Co-Reference Resolution, Relation Extraction or Event Extraction are encompassed under the umbrella of IE [2]. IE encounters specific challenges, particularly with regard to data availability and the need for expert knowledge. First, access to raw data is limited depending on the target domain (e.g.
Studying PH variability in coastal areas using deep learning - Actu IA
Seawater has a pH of about 8.2, although it can vary between 7.5 and 8.5 depending on local salinity, and is estimated to have declined on average by 0.1 since the industrial era. This downward trend associated with increasing CO2 levels in the atmosphere is a matter of concern because of the possible negative consequences for marine organisms, especially calcifiers (corals, shellfish …). A team of Spanish researchers conducted a study to assess the seasonal variability of pH. Entitled " pH trends and seasonal cycle in the coastal Balearic Sea reconstructed through machine learning", it was published in the journal Natureon July 28. Susana Flecha, Àlex Giménez-Romero, Joaquín Tintoré, Fiz F. Pérez, Iris E. Hendriks, Manuel A. Matías, Eva Alou-Font are the authors of this study, which aims to study the variability of the PH of the Balearic coastal area through deep learning.
HEROHE - Grand Challenge
Unlike previous Challenges that evaluated the staining patterns present in IHC, this Grand Challenge new edition proposes to find an image analysis algorithm to identify with high sensitivity and specificity HER2 positive BC from HER2 negative BC specimens evaluating only the morphological features present on the hematoxylin and eosin (HE) slide. The team as a whole has experience in machine learning and computer vision, has medical expertise, and has experience in organizing previous challenges. For questions related to the Grand Challenge, please contact Eduardo Conde-Sousa (econdesousa@gmail.com). This challenge is held as part of the European Congress on Digital Pathology. Please read the Rules section.
Instituto de Astrofísica e Ciências do Espaço
Astrophysicists use artificial intelligence to determine exoplanets sizes 2019 October 09 This artist's impression shows several of the planets orbiting the ultra-cool red dwarf star TRAPPIST-1. KornmesserTrue radii as a function of the predicted radii for the test set. Credit: Ulmer-Moll et al.A team1 of Instituto de Astrofísica e Ciências do Espaço (IA2) researchers has published an article3, led by Solène Ulmer-Moll, which shows that by knowing an exoplanet's mass and equilibrium temperature, it's possible to constrain its radius, with higher accuracy than previous methods. Solène Ulmer-Moll, a PhD student at the Science Faculty of the University of Porto (FCUP) explains this result was obtained by using knowledge from different fields: "This novel way to forecast exoplanet radius is a perfect example of the synergy between exoplanet science and machine learning techniques." To characterize a planet, both its mass and radius are needed, in order to find the planet's density, and from that infer its composition.
Predicting assisted ventilation in Amyotrophic Lateral Sclerosis using a mixture of experts and conformal predictors
Pereira, Telma, Pires, Sofia, Gromicho, Marta, Pinto, Susana, de Carvalho, Mamede, Madeira, Sara C.
Amyotrophic Lateral Sclerosis (ALS) is a neurodegenerative disease characterized by a rapid motor decline, leading to respiratory failure and subsequently to death. In this context, researchers have sought for models to automatically predict disease progression to assisted ventilation in ALS patients. However, the clinical translation of such models is limited by the lack of insight 1) on the risk of error for predictions at patient-level, and 2) on the most adequate time to administer the non-invasive ventilation. To address these issues, we combine Conformal Prediction (a machine learning framework that complements predictions with confidence measures) and a mixture experts into a prognostic model which not only predicts whether an ALS patient will suffer from respiratory insufficiency but also the most likely time window of occurrence, at a given reliability level. Promising results were obtained, with near 80% of predictions being correctly identified.