Cross-kingdom phenotype annotations for 35,856 species generated with a web-search-enabled language model

This article has 0 evaluations Published on
Read the full article Related papers
This article on Sciety

Abstract

Standardized organismal trait matrices support comparative biology, macroecology and conservation, but broad taxonomic coverage is difficult to obtain by manual curation or corpus-bound extraction. This Data Descriptor presents a cross-kingdom phenotype dataset generated by a web-search-enabled large language model embedded in a scripted annotation pipeline. The dataset contains 7,027,776 ordinal phenotype annotations for 35,856 animal, plant and fungal species across 196 questions spanning cellular biology, anatomy, physiology, ecology, behavior, life history, reproduction, environmental tolerance and human associations. Each annotation uses a five-level scale from Definitely No to Definitely Yes, including an Unknown or not-applicable category. The repository provides organism and question metadata, raw and processed annotation matrices, imputation outputs, derived phenotype scores, validation tables and analysis code. Technical validation includes comparisons with independent trait databases, a focused audit of 200 organisms, latent-structure analyses and phenotype-similarity checks. The data are intended for exploratory comparative analyses, hypothesis generation, benchmarking and targeted expert review, rather than as substitutes for primary measurements or curated species-level observations.

Related articles

Related articles are currently not available for this article.