Building a Governed Clinical Data Capability for a Whole Person Care Intelligence Company

A Whole Person Care Intelligence company partnered with ConceptVines to create clinically coherent synthetic patient populations for Type 2 diabetes and chronic obstructive pulmonary disease (COPD). The result was more than a large dataset: it was a reusable System of Work for designing, generating, reviewing, and delivering longitudinal clinical cohorts without exposing a single real patient record.

By combining generative AI with deterministic clinical rules, recognised healthcare standards, and expert review, ConceptVines produced more than 4,000 synthetic patients, over 7 million clinical data rows, and more than 10 years of history per patient. The platform gave research teams data they could use in existing analytics workflows while establishing an extensible foundation for additional disease areas.

Impact at scale

4,000+

Synthetic patients across Type 2 diabetes and COPD cohorts

7M+

Clinical data rows generated

10+ years

Longitudinal history per patient

Zero

Real patient records exposed in the delivered datasets

Client snapshot

The client is a Whole Person Care Intelligence company working to turn fragmented health information into a more continuous view of individual health, risk, and wellbeing. Its research teams needed substantial longitudinal datasets to develop and validate analytics models for two complex chronic-disease populations: Type 2 diabetes and COPD.

The data had to be useful from the outset. Diagnoses, disease severity, observations, medications, encounters, and comorbidities needed to form credible patient histories. Records also had to use established clinical terminology so they could move into the client’s analytics environment without a separate translation exercise.

Using real patient records would have introduced access, privacy, and governance dependencies before model work could begin. Conventional synthetic-data tools removed that dependency, but often created another problem: they could generate records at scale without reproducing the clinical logic required for meaningful longitudinal research.

The client did not simply need more data. It needed a dependable way to produce research-ready evidence.

The challenge

Synthetic healthcare data is often assessed by record count or statistical similarity. Neither measure is sufficient on its own.

A patient history is an interdependent sequence. COPD severity should influence pulmonary-function results, oxygen saturation, symptoms, medications, and subsequent encounters. Diabetes progression should be reflected in HbA1c trends, related observations, treatment patterns, and plausible comorbidities. Events must occur in a credible order, use the correct units and codes, and stop at the appropriate point in the patient timeline.

When fields are generated independently, a dataset can appear complete while contradicting itself. Stable patients may show implausibly abnormal observations. Severe conditions may have no corresponding treatment pattern. Values may exceed physiological limits, and clinical events may continue after a recorded date of death. These are not cosmetic defects; they compromise the validity of the research performed with the data.

The knowledge required to prevent these failures existed across disease models, terminology systems, physiological limits, medication logic, technical schemas, and the judgement of clinical reviewers. A generator alone could not coordinate them.

ConceptVines’ approach

ConceptVines designed the engagement around the work of cohort creation rather than around a single model or generation tool. The resulting platform brings clinical intent, terminology standards, deterministic rules, generative capability, longitudinal scheduling, automated validation, expert review, and downstream delivery into one governed process.

A shared representation of each patient

Each synthetic patient is modelled as a connected clinical history. The patient’s condition set and severity act as upstream signals shaping diagnoses, symptoms, comorbidities, observations, medications, and encounter patterns.

This preserves relationships across the record. COPD severity can influence pulmonary-function measures and treatment choices. Diabetes progression can influence HbA1c trends and associated observations. Chronic kidney disease staging can progress logically rather than shifting arbitrarily between encounters.

Generative variation within deterministic boundaries

Generative AI creates diverse, schema-constrained histories. Deterministic rules control the elements for which inconsistency would reduce research value or introduce clinical error, including physiological ranges, disease-severity relationships, medication conflicts, timeline integrity, and referential links across clinical entities.

The model is therefore not treated as the authority on clinical truth. It operates inside an architecture that assigns probabilistic intelligence and deterministic control to the work each is suited to perform.

Longitudinal, standards-native data

A rule-based scheduler creates outpatient, inpatient, and emergency encounters across more than a decade. Patient-level state is preserved so observations change over time rather than being regenerated without memory. When reviewers found that four years of history was insufficient for trend analysis, ConceptVines extended the simulation window to more than 10 years.

Diagnoses use ICD-11, observations use LOINC, and medications use RxNorm. By embedding these standards during generation, the platform produces data that can enter the client’s existing analytics pipeline without an additional mapping layer.

Clinical judgement inside the system

The client’s clinical team reviewed successive samples against reference parameters for diabetes and COPD. ConceptVines converted their findings into new rules, fields, constraints, and tests, typically returning revised samples within days.

Missing COPD measures were added with the appropriate codes. SpO2 values above 100% and incorrectly represented FEV1/FVC ratios were corrected through hard constraints and testing. Severity-scaled ranges were introduced when stable patients displayed abnormal values too frequently.

Human review was not a final approval gate applied after generation. It became part of the system’s learning loop, strengthening the reusable clinical context governing every subsequent cohort.

The impact

The platform produced 2,500 synthetic diabetes patients and scaled the COPD population from a 150-patient validation sample to a 1,500-patient cohort. Across the delivered populations, it generated more than 7 million rows of clinical data.

The diabetes cohort alone included approximately 3.8 million laboratory, vital-sign, examination, and social-determinants observations, alongside extensive diagnosis and medication histories. The value, however, extended beyond volume.

Research without real-patient exposure
Teams could begin analytics and model-development work without using real patient records. Synthetic data does not replace validation against appropriate real-world evidence before clinical use, but it provides a governed environment for developing pipelines, testing hypotheses, and finding model weaknesses earlier.
Longitudinal analysis
More than 10 years of history per patient enabled analysis of disease progression and trends rather than limiting research to point-in-time relationships.
Quality that scaled with volume
Automated checks covered physiological plausibility, terminology conformance, referential integrity, and longitudinal coherence as the cohorts moved from expert-reviewed samples to production scale.
Immediate integration
Standards-native coding allowed the datasets to load directly into the client’s clinical analytics pipeline, enabling researchers to focus on clinical realism and model behaviour rather than data-format remediation.
A reusable enterprise capability
The same underlying engine supports diabetes and COPD and can expand into additional chronic diseases through new condition, severity, and validation rules. The client now owns a capability that can accumulate clinical context, reviewer judgement, and quality controls over time.

What’s next

The next phase is intended to consolidate the diabetes and COPD pipelines behind a single configurable engine, introduce automated regression checks for each delivery, and extend the framework to additional chronic-disease populations.

This is how progressive autonomy develops in enterprise work. As the system captures more disease logic and validation knowledge, it can automate a greater share of cohort configuration and quality assurance while continuing to escalate clinically consequential questions to human experts.

Why ConceptVines

ConceptVines combined enterprise architecture, AI engineering, standards-aware data modelling, deterministic controls, and iterative clinical review in one operating system for the work.

A model could generate records, and a conventional pipeline could move them. Neither could independently coordinate the clinical relationships, terminology standards, validation logic, expert judgement, and longitudinal state required to produce dependable cohorts at scale.

For this Whole Person Care Intelligence company, the result was not simply more synthetic data. It was a governed System of Work capable of turning clinical intent into coherent longitudinal evidence, creating immediate research value while building an enterprise capability that can compound across new diseases, new models, and new questions.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top