Main discussion
In this study, we evaluated the prediction of treatment outcomes in OCD patients using a combination of demographic, clinical, and neuroimaging data through various ML models. We implemented a rigorous validation scheme to rule out data leakage, compared multiple ML algorithms to mitigate bias, and performed permutation testing to identify replicable treatment success markers. Treatment outcome was operationalized using two commonly employed binary labels: response, defined as a ≥ 35% reduction in Y-BOCS scores, and remission, defined as a post-treatment Y-BOCS score below the clinical threshold.
Despite extensive analyses in a well-characterized clinical sample, neither neuroimaging features alone nor multimodal models integrating neuroimaging with demographic and clinical variables produced prediction accuracies that exceeded chance level for either outcome criterion. Clinical and demographic features showed modest descriptive predictive value, particularly for remission prediction, where accuracies approached 66%, but these effects did not survive permutation testing. In contrast, prediction performance for response remained consistently at chance. Taken together, these findings suggest that robust individualized prediction of treatment outcome in OCD was not detectable in the present sample and modeling framework.
Although several clinical variables showed partial convergence with previous literature, no single feature or combination of features emerged as a consistently detectable predictors of treatment success. Prior ML studies similarly identified baseline OCD severity, depressive symptoms, and psychosocial functioning as modest but inconsistent predictors of remission5,10,11,12. Hilbert et al.11 achieved 65% balanced accuracy in predicting remission using baseline OCD severity, socioeconomic status, age, and depression scores, while Askland et al.10 achieved 75% balanced accuracy in a longitudinal sample, identifying baseline compulsion severity, social adjustment, and neuroticism as primary predictors. In the present study, SHAP-based feature importance rankings showed overlap with these findings. While we caution that these rankings are strictly descriptive and do not represent clinically meaningful predictors due to at-chance model performance, OCI-R symptom severity, depressive symptoms, socioeconomic status, and baseline Y-BOCS scores were among the most consistently ranked predictors across clinical models. Notably, baseline Y-BOCS severity appeared substantially more relevant for remission prediction than for response prediction.
A recognized fundamental challenge in extracting a consistent predictive signal is the clinical heterogeneity of OCD itself. As highlighted in a recent systematic review49, OCD is not a uniform disorder but a cluster of different phenotypes. Consequently, identical aggregate Y-BOCS scores may mask divergent symptom dimensions, etiologies, and cognitive profiles. This heterogeneity introduces substantial variance into the modeling process, making it significantly harder for supervised algorithms to identify generalizable patterns from baseline features.
Consistent with this complexity, the baseline clinical profiles of remitters and non-remitters showed differences across a range of measures. Specifically, non-remitters scored higher on measures of OCD symptom severity (washing, incompleteness), depressive symptoms, trait anxiety, neuroticism, shyness, anticipatory worry, and executive dysfunction. Rather than indicating distinct psychophenotypic subtypes, these differences are most consistent with a generalized severity and chronicity continuum in which patients with greater overall psychopathology and functional impairment are less likely to achieve favorable outcomes. For instance, elevated neuroticism and trait anxiety are well-established transdiagnostic markers of poorer therapeutic outcomes50,51. Furthermore, as highlighted in a recent review49, executive dysfunction may act as a specific barrier to the cognitive flexibility and active engagement required for successful CBT.
This severity dimension has direct implications for how remission-specific prediction results should be interpreted. Because remission is defined by an absolute post-treatment symptom threshold, patients entering treatment with lower baseline severity require a smaller absolute symptom reduction to achieve remission. Consequently, models assigning strong weight to baseline Y-BOCS scores may partly recover proximity to the remission threshold rather than individualized treatment prognosis. Remission prediction should therefore be interpreted cautiously as reflecting the likelihood of reaching a low-symptom endpoint rather than purely predicting treatment-related change. This interpretation is consistent with the observation that baseline severity contributed more strongly to remission prediction than to response prediction. In contrast, the response criterion is based on proportional symptom reduction and is therefore substantially less dependent on initial symptom severity.
Given the distinct limitations of each criterion when applied in isolation, we retained both response and remission as parallel outcome measures. The response criterion captures clinically meaningful proportional symptom improvement but may represent an overly stringent bar for patients entering with mild-to-moderate baseline severity, while the remission criterion reflects a clinically meaningful absolute symptom endpoint but is susceptible to the threshold proximity effects described above. Reporting both criteria in parallel partially offsets these complementary shortcomings and allows the degree of convergence between them to serve as an additional interpretive lens.
A counterintuitive pattern in the data illustrates this difference concretely: seven patients achieved remission without meeting the response criterion. This subgroup consists of patients who entered the study near the lower bound of the inclusion criterion, for whom a modest absolute improvement was sufficient to achieve remission but fell just short of the reduction threshold required for response. Rather than a data anomaly, this subgroup is a direct consequence of combining an absolute with a relative criterion.
Beyond the choice between outcome criteria, these considerations also bear on the decision to use binary rather than continuous outcomes. The practical goal is to identify patients at elevated risk of non-response so that augmented or alternative treatment pathways can be considered from the outset. A binary classification maps directly onto this decision, while a regression model’s output would still require a threshold to translate into an actionable clinical recommendation, a step that is no easier to calibrate and that does not circumvent the underlying classification problem. To fully address the potential inflation of baseline influence due to threshold dependency, we conducted a supplementary regression analysis predicting percent symptom change as a continuous outcome (see Supplementary Material). Despite being theoretically more powerful and less susceptible to threshold effects, this continuous approach likewise did not yield a detectable predictive signal. The consistency of findings across both binary and continuous operationalizations suggests that the lack of a predictive signal is unlikely to be explained solely by outcome dichotomization.
Turning to neuroimaging predictors, neither structural nor resting-state fMRI features yielded above-chance prediction, despite a systematic and progressive approach to feature representation. Previous studies examining fMRI network connectivity in OCD have not yet identified a coherent set of predictive features19,20,21,22. Therefore, we evaluated multiple levels of feature abstraction in sequence, including whole-brain functional connectivity matrices with regularization-based dimensionality control, PCA/ICA-based feature reduction, and compact global graph measures summarizing large-scale network topology. While necessary to manage our sample size, whole-brain PCA/ICA risks discarding weak localized signals. To ensure that dimensionality reduction did not obscure critical information, we conducted a supplementary hypothesis-driven analysis using unreduced amygdala–vmPFC connectivity19. This targeted approach also yielded chance-level predictions (see Supplementary Material), supporting the conclusion that the findings are not specific to PCA-based feature reduction. That above-chance prediction was not achieved at any level of feature abstraction may reflect the limited sample size available for imaging analyses as much as a genuine absence of relevant signal in the functional connectivity data. Node-level graph metrics were not included in the present analyses but represent a meaningful extension for future work, e.g. applied to frontolimbic cicuits19 or default-mode network nodes21. In line with a previous large-scale study26, structural MRI features similarly did not contribute to prediction.
The difficulty in establishing individualized brain-based predictions aligns with recent literature discussing the feasibility of finding associations between psychological phenotypes such as clinical presentation and imaging-derived markers. Large-scale studies52,53 suggest that sample sizes in the range of hundreds or thousands of patients might be needed for this task. Two recent treatment outcome prediction studies in mental disorders54,55 examined generalization performance of their ML models on independent clinical cohorts. Classification accuracies in the studies exceeded chance level but were rather low overall (58% in 55; 56% − 67%, in 54) and the models failed to generalize to external datasets. Ruffle and colleagues56 conducted a systematic evaluation of the predictability of individual characteristics from structural and fMRI data and reported poor prediction performance for psychological characteristics. In the majority of cases, models based on neuroimaging data did not contribute to prediction performance of participants’ mental health variables obtained from other health-related features.
A review on treatment outcome prediction in mental health interventions by Vieira et al.9 finds the field to be subject to publication bias with an over-reporting of larger effects. These stronger results stem from studies with small sample sizes and are not matched by studies demonstrating the same effects at larger scales. Supporting the conclusion by Vieira and colleagues9, several previous studies have employed less rigorous or no CV schemes and no permutation testing, increasing the risk of inflated effect sizes. This particularly afflicts neuroimaging studies that typically examine smaller sample sizes. In combination with overreporting of positive results this may have painted an overoptimistic picture of the prediction accuracies achievable for current ML models in precision psychotherapy.
In this light, our findings are consistent with the broader challenge of predicting psychological outcomes, potentially due in part to label noise57, even in relatively large samples. This challenge is likely to be even greater for neuroimaging studies, which generally require substantially larger sample sizes to reliably detect predictive signals.
Limitations
Our dataset allowed us to perform CV; however, its relatively small sample size limited the power of our ML analyses (see Supplementary Analysis 3). For neuroimaging analyses, particularly fMRI, the reduced sample size may have been insufficient to achieve above-chance predictive performance. The dataset also lacked sufficient power to detect statistically significant effects in models descriptively performing above-chance. For example, although analyses including the established predictor symptom severity showed classification accuracies of up to 66% for remission, the sample size might not have provided adequate power to confirm these results statistically.
The neuroimaging analyses were subject to a particularly challenging dimensionality-to-sample-size ratio. Structural MRI yielded 496 features, and resting-state functional connectivity matrices span a substantially larger feature space prior to dimensionality reduction. Several methodological choices were made to mitigate this problem. First, we used nested CV to obtain unbiased performance estimates. For dimensionality reduction of functional connectivity features, we applied PCA/ICA. We further used regularized classifiers with penalization tuned in the inner loop, and variance thresholding to remove uninformative features prior to modelling. Finally, we used global graph summary measures rather than node-level metrics to keep the feature space compact. These steps reduce overfitting risk but cannot compensate for a fundamental limitation of statistical power. The null imaging results may therefore reflect the constraints of the current sample and design at least as much as a genuine absence of predictive signal. Replication in larger, adequately powered cohorts would help clarify the utility of neuroimaging features for treatment outcome prediction in OCD.
Better clinical prediction may be achieved by considering more fine-grained input data. Our clinical feature set included a variety of questionnaires assessing OCD symptomatology, comorbidities, and personality traits. However, we only used summed subscale scores from these instruments. Previous work by Askland et al.10 and Hilbert et al.11 suggests that individual items measuring OCD symptomatology may constitute predictors beyond overall symptom severity. Item-level features may therefore provide valuable additional information for prediction analyses. Future studies should exploit this source of information by focusing on deep phenotyping involving specific sets of relevant symptoms that are interrogated at a fine-grained level rather than a broad set of questionnaires covering many areas of impairment. With such fine-grained data, future research could also employ data-driven clustering approaches to empirically determine whether complex clinical profiles represent discrete etiological phenotypes or simply endpoints along a shared severity dimension.
Furthermore, our dataset did not include detailed characteristics of the therapeutic process itself. While our models were specifically designed to utilize pre-treatment data to inform decisions at the outset of therapy, CBT-related process variables, such as specific protocol components, therapeutic alliance ratings, homework adherence, therapist fidelity scores, and session-by-session symptom change trajectories, are critical factors in determining treatment success. Future studies that prospectively record these process variables alongside comprehensive baseline assessments are highly encouraged. Integrating these factors into ML frameworks could capture mechanisms of change that baseline assessments cannot, potentially yielding stronger prediction models than those achievable from pre-treatment data alone.
We acknowledge as a limitation our use of a last-observation-carried-forward approach for the 26 patients who discontinued therapy prematurely. While this strategy can potentially bias outcomes, our data suggests that premature discontinuation did not systematically correspond to poor treatment outcome. Relatedly, the substantial heterogeneity in treatment dose and duration across our sample may serve as a source of label noise. However, our supplementary analyses offer reassurance that session count itself does not carry a strong, systematic predictive signal.
Finally, we compared the outcome labels response and remission, which are subject to considerable label noise since they were based on Y-BOCS scores alone. Additional related clinical or real-life measurements may improve the quality of those labels. For example, quality of life measures, social functioning, or outcome measures that are based on longer time intervals may potentially be more robust, as seen in Askland et al.10 who used longitudinal measures assessing longer time ranges to determine the criterion (percent time spent in remission and ever experiencing remission that lasted at least 8 weeks). Using such a criterion, treatment outcome was predictable with 75% accuracy. Future studies with larger samples should explore continuous outcome modeling as a complementary approach, as such analyses would be free of the threshold dependency introduced by binary remission criteria.
Conclusion
In this study, we used a combination of demographic, clinical, and neuroimaging data to predict treatment outcome in OCD patients following CBT. While clinical predictors descriptively performed better than neuroimaging features, no statistically significant predictive signal was detectable in the present sample, and integrating neuroimaging data did not enhance prediction accuracy beyond demographic and clinical variables alone. These findings align with broader literature highlighting the limitations of current neuroimaging techniques for individualized prediction in mental health but need to be cautiously interpreted given their feature dimensionality relative to the sample size. The modest prediction accuracies underscore the need for larger samples, as well as more fine-grained input features and more reliable outcome criteria. Future research should focus on deep phenotyping and longitudinal studies to enhance the robustness and applicability of predictive models.
