LongevityBench: An Open Benchmark and Language Models for AI in Aging Biology (Liquid AI / InSilico Medicine)
- Type
- paper
- Venue
- Liquid AI / InSilico Medicine; Cell, vol. 189, no. 19, Sept. 2026, pp. 5980-5994.e8
- Year
- 2026
- Source
- blog
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
LongevityBench enables systematic evaluation of language models on aging- and longevity-related data across five biodata domains: clinical records from NHANES, DNA methylation studies from GEO, transcriptomic profiles from GTEx, plasma proteomics from three Olink studies, and genetic evidence from OpenGenes, CellAge, and SynergyAge. Comprises 17 tasks and 25,457 prompts in four formats: binary classification, pairwise comparison, multiclass classification, and numeric regression. Tasks are separated from training by biology/study-level splits (NHANES by survey wave, methylation by GEO study, transcriptomics/proteomics by individual, OpenGenes by protein family, SynergyAge by species) for leakage-proof evaluation. The accompanying domain-adapted LFM2-1.2B-Longevity and LFM2-2.6B-Longevity (LFM2-1.2B/2.6B fine-tuned 3 epochs, 32k context, on a multitask aging-prompt collection plus 10-20% general chat to preserve conversational ability, merged by equal-weight linear combination) often matched or exceeded much larger frontier models compared against 18 frontier LLMs: LFM2-2.6B-Longevity ranked first on the NHANES pairwise-age task, both LFMs topped the masked OpenGenes gene-expression-direction task, second and fourth on GTEx multiclass age-group (both beating every frontier model), second on GEO pairwise-age, and both outperformed every frontier model on the Olink proteomics pairwise-age task. A matched-ablation analysis maps which biological feature groups drive predictions. Compact models enable local deployment on sensitive patient data. Benchmark, models, and paper released publicly.
Keywords
benchmarks · aging · longevity · biology · domain adaptation · multimodal structured data
Topics
biology, aging, longevity, benchmarks, domain adaptation
Research notes
- Discovery: Liquid AI (@liquidai, verified) 2026-09-30: https://x.com/liquidai/status/2105311716976521347 (with @InSilicoMeds)
- Published paper: Zhavoronkov et al., "An Open Benchmark and Language Models for AI in Aging Biology," Cell: https://www.cell.com/cell/fulltext/S0092-8674(26)00999-2
- Hugging Face: LongevityBench benchmark plus LFM2-1.2B-Longevity and LFM2-2.6B-Longevity.
- Presented at the ARDD Meeting (aging research) the week of the announcement.
- Connects to the collection's benchmarks, biology, domain-adaptation, and leakage-proof-evaluation entries.