The ProPark Data Platform is a secure, informed consent governed infrastructure that turns the multimodal data of the ProPark observational clinical study into analysis-ready datasets for Parkinson's disease research. It handles clinical examinations, questionnaires, medication records, and continuous wearable sensor recordings, with the architecture to expand into any data type in the future.
ProPark (Profiling Parkinson's Disease) is a multi-year observational study led by Leiden University Medical Center with participating sites across the Netherlands, following approximately 900 people with Parkinson's disease alongside healthy controls. Beyond clinic-based motor examinations and questionnaires, participants wear motion sensors on the wrist and lower back during repeated week-long home assessments, which captures how the disease actually presents across medication states and ordinary daily activity, rather than only as it appears in a brief clinic visit.
Data of this scale and variety creates a corresponding infrastructure challenge: trillions of sensor measurements, heterogeneous clinical records, and the practical messiness that is normal in real clinical research. Since 2023, OccamzRazor has served as the study's data engineering partner. We design, build, and operate the platform that receives the data. The platform organizes and cleans it, protects it, and gives researchers a controlled and user-friendly way to analyze it.
Data moves through five governed stages, from raw study files to research-ready analysis.
Sensor files, clinical exports, and medication forms are received into dedicated secure storage configured so that uploaded files cannot be altered or deleted after submission — preserving an unmodified record of what was received.
Incoming data is validated, cleaned, and standardized. Wrist and lower-back recordings are time-synchronized to a common clock, and validated algorithms extract clinically meaningful measures — tremor, bradykinesia, gait, and overall activity — with thousands of processing jobs running in parallel.
Processed data is organized into linked sensor, feature, and clinical tables, so that measures collected for the same participant at the same point in the study can be brought together into a single analysis dataset.
Researchers query the data on demand in standard SQL — retrieving exactly the participants and time periods relevant to an analysis, without operating any infrastructure of their own. Analytical queries that once took days or weeks of dataset preparation now run in seconds.
Where analysis requires expert judgment — such as labeling motor tasks within a recording — data segments are routed through a structured, secure annotation workflow, and the resulting labels are returned to the database without ever leaving the governed environment.
Research infrastructure for patient data is only as trustworthy as its weakest control. The platform was built from the ground up so that its central commitments do not depend on any individual applying the rules correctly.
The platform maintains a record of each participant's data-sharing permissions. Researchers are assigned to defined categories and query controlled views that return only the participants whose consent permits that category of use — enforced automatically by the platform itself.
All primary storage and processing take place within a European Union cloud region, consistent with the GDPR and with the privacy commitments made to study participants.
Data is encrypted at rest and in transit. Ingestion pathways are versioned and deliberately constrained, and administrative access runs through controlled, recordable channels.
Every transformation applied to a clinical value is recorded in a transformation log, and access to data is logged. Any result can be traced back to the originally recorded value, and any access can be reviewed.
We asked, what would it take for a neurologist to explore a hypothesis as soon as the idea occurs to them? Then we built the systems to remove the analysis bottlenecks, dramatically increasing the pace of discovery.
Cohort-wide tremor feature extraction, previously an ad-hoc workflow across local machines, now runs in the platform's parallel compute environment in a fraction of the time.
Creating a study dataset once took days to weeks of manual assembly. Researchers now pose analytical queries directly against the governed data store and receive answers interactively.
Because the data is clean, documented, and governed, AI-assisted analysis can operate directly on it. In a live demonstration for Dutch research funders, a published peer-reviewed analysis from the study was reproduced end-to-end — figures included — through plain-language interaction with the data.
The platform's purpose is published science. Its data engineering underpins peer-reviewed research from the ProPark study, and its outputs are built to be traceable from figure back to raw signal.
In 219 ProPark participants, wearable-derived tremor measures showed excellent test–retest reliability and agreed with clinical assessment — and detected tremor that patients reported but brief clinic exams missed. The platform ingested and processed the raw sensor data behind the study and integrated it with clinical data to enable the analyses.
OccamzRazor is the data and modeling team behind neurodegeneration research. We build infrastructure and tooling that turns complex datasets into the technologies that empower researchers to transform the lives of patients. Our partnership with LUMC demonstrates how we work: start with utmost respect for patient privacy, embed ourselves inside the research teams tackling the world's most important problems in neurodegeneration, and remove the barriers between their questions and the answers.