BEND: Benchmarking DNA Language Models on biologically meaningful tasks
Marin, Frederikke Isa, Teufel, Felix, Horlacher, Marc, Madsen, Dennis, Pultz, Dennis, Winther, Ole, Boomsma, Wouter
–arXiv.org Artificial Intelligence
The genome sequence contains the blueprint for governing cellular processes. While the availability of genomes has vastly increased over the last decades, experimental annotation of the various functional, non-coding and regulatory elements encoded in the DNA sequence remains both expensive and challenging. This has sparked interest in unsupervised language modeling of genomic DNA, a paradigm that has seen great success for protein sequence data. Although various DNA language models have been proposed, evaluation tasks often differ between individual works, and might not fully recapitulate the fundamental challenges of genome annotation, including the length, scale and sparsity of the data. In this study, we introduce BEND, a Benchmark for DNA language models, featuring a collection of realistic and biologically meaningful downstream tasks defined on the human genome. We find that embeddings from current DNA LMs can approach performance of expert methods on some tasks, but only capture limited information about long-range features. BEND is available at https://github.com/frederikkemarin/BEND.
arXiv.org Artificial Intelligence
Nov-25-2023
- Country:
- Africa
- Kenya (0.04)
- Nigeria > Oyo State
- Ibadan (0.04)
- Sierra Leone (0.04)
- The Gambia (0.04)
- Asia
- Bangladesh (0.04)
- China
- Beijing > Beijing (0.04)
- Guangdong Province > Shenzhen (0.04)
- Japan > Honshū
- Kantō > Tokyo Metropolis Prefecture > Tokyo (0.14)
- Middle East > Israel
- Tel Aviv District > Tel Aviv (0.04)
- Pakistan > Punjab
- Lahore Division > Lahore (0.04)
- Vietnam > Hồ Chí Minh City
- Hồ Chí Minh City (0.04)
- Europe
- Denmark > Capital Region
- Copenhagen (0.04)
- Finland (0.04)
- Germany > Bavaria
- Upper Bavaria > Munich (0.04)
- Italy (0.04)
- Netherlands > South Holland
- Leiden (0.04)
- Spain (0.04)
- United Kingdom
- England > Oxfordshire
- Oxford (0.14)
- Scotland (0.04)
- England > Oxfordshire
- Denmark > Capital Region
- North America
- Barbados (0.04)
- Canada > Quebec
- Montreal (0.04)
- Cuba > La Habana Province
- Havana (0.04)
- Puerto Rico (0.04)
- United States
- California
- Los Angeles County > Los Angeles (0.14)
- San Diego County > San Diego (0.04)
- San Francisco County > San Francisco (0.14)
- Santa Cruz County > Santa Cruz (0.04)
- Illinois > Cook County
- Chicago (0.04)
- Northbrook (0.04)
- Massachusetts (0.04)
- Pennsylvania (0.04)
- Texas > Harris County
- Houston (0.14)
- Utah (0.04)
- California
- South America
- Chile > Santiago Metropolitan Region
- Santiago Province > Santiago (0.04)
- Colombia > Antioquia Department
- Medellín (0.04)
- Peru > Lima Department
- Lima Province > Lima (0.04)
- Chile > Santiago Metropolitan Region
- Africa
- Genre:
- Research Report > New Finding (0.48)
- Industry: