Community Survey: Does data standardization still matter in the era of foundation models?

Hi all,

I’m Jiwon Um, a PhD candidate at Yonsei University and OHDSI APAC Community Call Lead. I’d like to share a community survey our group has launched, and to invite your participation.

The question

Foundation models — LLMs, multimodal systems — are reshaping how biomedical data is processed. They can read unstructured notes directly, work with imaging and waveforms, and adapt to schema-flexible inputs.

This creates a genuine tension. Schema-flexible learning prompts some to question whether structured ETL remains a favorable cost-benefit ratio, while federated validation and hallucination control appear to require, rather than supersede, a common semantic layer.

So a real question is emerges:

Is OHDSI’s decade-long investment in OMOP CDM, vocabularies, and federated networks still a trust mechanism for AI evidence — or is it becoming a legacy cost?

While discussions about foundation models and standardization are growing across medical informatics, the community itself has not been systematically asked where it stands.

The gap

Recent surveys have studied clinicians’ attitudes toward AI (e.g., AMA 2026, n=1,692) and global expert assessments of AI failure modes (Müller et al., 2026, n=914).

But the technical community that builds and maintains the data infrastructure underlying medical AI — that’s us — has never been similarly studied. This survey aims to fill that gap.

Join us

The protocol, instrument, and analysis plan are pre-registered and public before any data is collected — you can review everything before participating.

I’d glad to answer any questions and for keeping this community what it is.

— Jiwon Um
PhD Candidate, Biomedical Systems Informatics, Yonsei University