LLMs extract clinical data from unstructured EHR text
A disease-agnostic framework using large pre-trained language models can extract computable clinical data from unstructured electronic health record text, according to a study published in Nature on Sept. 10, 2026.
Researchers presented a novel approach using large, pre-trained language models to accurately extract computable clinical data from unstructured text in electronic health records, according to a study published in Nature on Sept. 10, 2026.[1]
The framework is disease-agnostic and scales to large patient datasets across all clinical conditions.[2] Extracted clinical entities are integrated with structured EHR data, embedded with medical ontologies, and organized into a knowledge graph allowing relationships to be interrogated at the level of individual patients or at scale.[3]
The read here is that the approach could reduce manual chart review for dietitians and clinicians who rely on structured data from EHRs for malnutrition screening, nutrition assessment, and monitoring of enteral or parenteral nutrition outcomes. By converting free-text notes into computable data linked to ontologies, the framework may support clinical decision support and population-level nutrition research, though the study did not report specific performance metrics or validation against chart review.
In related work, researchers proposed a decision support system using large language model features to recommend chest CT protocols from free-text clinical indications to address inconsistencies in manual selection, reported on arXiv.org on Sept. 9, 2026.[4] Another group presented MMTClinic, a benchmark designed to evaluate large language models on complex reasoning and question-answering tasks involving clinical time-series data, reported on arXiv.org on Sept. 7, 2026.[5]
Proof5 sources · 2 publishers · signed
What this stands on
Researchers presented a novel approach using large, pre-trained language models to accurately extract computable clinical data from unstructured text in electronic health records. · Nature
The framework is disease-agnostic and scales to large patient datasets across all clinical conditions. · Nature
Extracted clinical entities are integrated with structured EHR data, embedded with medical ontologies, and organized into a knowledge graph allowing relationships to be interrogated at the level of individual patients or at scale. · Nature
Researchers proposed a decision support system using large language model features to recommend chest CT protocols from free-text clinical indications to address inconsistencies in manual selection. · arXiv.org
Researchers presented MMTClinic, a benchmark designed to evaluate large language models on complex reasoning and question-answering tasks involving clinical time-series data. · arXiv.org
We could not place any of them by their address. None is an official body: that part stands on reporting, not on the underlying document or transcript.