An interpretable model for repeated mobility data produced the strongest held-out predictions in a comparison spanning disease cohorts and five clinical outcomes. DeMMO's overall normalized mean squared error, or nMSE, was 0.769, compared with 0.787 for the strongest baseline. Since lower nMSE means less error, that represents better performance. Its weighted correlation, or wR, was 0.464 versus 0.448, where a higher score indicates a closer match between predictions and outcomes. The report gives averages over five matched participant splits, but no confidence intervals, p-values or named significance tests.
A test built around participants
The document is an arXiv version 2 preprint dated 30 Aug 2026. The evaluation divided participants into 70% training, 10% validation and 20% held-out test sets. All visits from each participant stayed in the same set, and the procedure was repeated over five matched splits. DeMMO was compared with eight baseline methods, with tuning settings selected using the validation data.
What went into the model
The study asked whether repeated multivariate digital mobility outcomes, or DMOs, could support joint prediction of multiple clinical outcomes across mobility-limiting diseases. After requiring a valid target and complete observations for all 24 weekly DMOs, the retained participant counts were 574 for the PD H&Y outcome, 574 for PD MDS-UPDRS III, 578 for multiple sclerosis EDSS, 469 for the PFF SPPB impairment outcome and 583 for COPD FEV1 percentage predicted.
To represent PFF impairment, the analysis transformed the SPPB score to 12 minus the SPPB score. It standardized each outcome and the 24 DMOs using pooled training records, then applied those training-derived settings unchanged to the validation and test data.
Sharing patterns, not people
At its core, DeMMO learns a symmetric signed relation graph from longitudinal coefficient mappings. In plain terms, it looks for mobility patterns that line up or point in opposite directions across outcomes, then shares selected information between cohorts whose participant records do not overlap. It does not pool participant records.
The advantage held across visits
The model's edge was not confined to the overall score. DeMMO had the lowest mean root mean squared error, or RMSE, in 21 of 25 outcome-visit tasks. It led at all five visits for PD H&Y and COPD FEV1, at visits two through five for PD MDS-UPDRS, at visits one through three and five for multiple sclerosis EDSS, and at visits three through five for PFF SPPB. Other methods led at the four remaining comparisons. These visit-level results were also averages over five matched splits, and no visit-level significance test was displayed.
The patterns were specific to each outcome
The learned relation weights offered a view of how the model shared information. The strongest positive relation was between the two PD outcomes, at 0.37, followed by the relation between multiple sclerosis EDSS and PFF SPPB, at 0.34. The strongest negative relations linked PD MDS-UPDRS with COPD FEV1 at negative 0.30 and with multiple sclerosis EDSS at negative 0.19. These values describe the orientation of model coefficients, not disease associations or correlations between raw clinical scores.
A longitudinal stability analysis highlighted two measures for multiple sclerosis EDSS: walking-speed P90 and cadence P90 in walking bouts longer than 30 seconds. Their selection probabilities were 0.878 and 0.874. Selection varied more across objectives than across visits, with reported variation of 0.111 versus 0.019. No DMO had a mean selection probability above 0.4 for all five objectives. The selected patterns are candidates for clinical validation, not validated biomarkers.
Exploratory disability tests exposed a weakness
The researchers also compared the original convex formulation with two non-convex variants. DeMMO recorded an overall nMSE of 0.769 plus or minus 0.021, compared with 0.773 plus or minus 0.035 for DeMMO-var1 and 0.772 plus or minus 0.033 for DeMMO-var2. It also had the lowest RMSE in 12 of 25 visit-level tasks, versus seven for DeMMO-var1 and six for DeMMO-var2.
In an exploratory binary EDSS analysis, a decision tree using two DMOs classified 2,083 participant-visit observations from 578 people. The observations included 1,157 non-severe and 926 severe cases. The tree achieved balanced accuracy of 0.789 plus or minus 0.031 and macro-F1 of 0.785 plus or minus 0.032. It correctly identified 84.2% of severe observations and 73.7% of non-severe observations, compared with a balanced accuracy of 0.5 for an always-most-common-class baseline. This was an observed-visit, retrospective classification analysis rather than prospective detection of deterioration.
The three-class version was less even. It used 317 mild, 840 moderate and 926 severe observations, producing balanced accuracy of 0.593 plus or minus 0.023, macro-precision of 0.543 plus or minus 0.016 and macro-F1 of 0.533 plus or minus 0.030. Recall was 70.8% for mild observations, 26.4% for moderate observations and 80.7% for severe observations. The authors describe this analysis as hypothesis-generating and note that it had poor resolution of the intermediate class.
A model comparison, not a clinical tool
The EDSS thresholds and classifications are associations learned from Mobilise-D and are hypothesis-generating rather than diagnostic cut-offs. The observed-visit experiments did not implement the complete software pipeline or prospectively detect individual deterioration. The mobility measures highlighted by the model therefore remain candidates for clinical validation.
The paper states that its model, optimization, data splitting, preprocessing, hyperparameter selection and evaluation protocol are specified. It also says that source code and privacy-safe intermediate and aggregate results are publicly available.
Paper data and sources
Original title: DeMMO: Longitudinal and Cross-Disease Modelling of Digital Mobility Outcomes via Multi-Task Learning
Authors: Menghui Zhou, Zhipeng Yuan, Vitaveska Lanfranchi, Po Yang
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text