Call for collaborators: detecting cross-site semantic divergence before network studies

I am looking for two sites willing to help validate a method. The ask is at the bottom if you want to skip the background.

Disclosure first. I am developing this work commercially. I am not selling anything here and there is no cost to participate. What I need is validation help, and this community is where the expertise sits.

The problem. When two sites both report a cohort of heart failure patients, they may not mean the same thing. Mapping choices, source vocabularies, and local documentation practice differ. The same concept can cover different patients at different places. That difference does not surface in a study result. It just makes the result quietly wrong.

Existing data quality tooling checks whether each site’s data is internally sound. What I am working on is different. It compares two sites against each other and asks whether a given concept carries the same clinical meaning at both.

The method. It builds a profile for each concept at each site from several signals, including co-occurring measurements, drug context, and specialty context. It then compares those profiles pairwise and flags concepts that diverge. The computation runs locally at each site and no patient level data moves.

I have only tested this against synthetic OMOP data where I injected the divergence myself. So I know it finds what I put there. I do not know whether it finds what actually exists in real data. That is the gap I am trying to close.

The ask. Two sites, because the method is a comparison and one site has nothing to compare against. The easiest version is one organization running it at two of its own instances, since that is a single conversation instead of two.

What running it involves is a read only package against your CDM. It writes nothing back. No patient level data leaves your environment. You see the output before I do and you control what is shared.

What I need beyond the run, is someone who knows your data well enough to look at the flagged concepts and tell me whether the differences are real or whether the method is wrong. Anyone can run code. What I need is the adjudication.

Happy to answer questions here or schedule time to chat. I will also be presenting at the 2026 Global Symposium in New Jersey in October if it is easier to talk in person.

Sandy Estremera-Zink