Abstract
Background conditions in clinical machine learning are now favorable for model development and still unfavorable for dependable translation. Hospitals capture large amounts of perioperative, physiologic, laboratory, and administrative information, allowing investigators to train increasingly sophisticated risk models. Yet these data are generated inside care systems that differ in surveillance intensity, intervention timing, documentation structure, and event ascertainment. A model that performs well within one such system may therefore encode a relationship that is only partly portable. This study examined that problem empirically in a multi-hospital cohort of major abdominal surgery patients and asked whether external validation, rather than source-domain optimization, determined the main limit on translational credibility. We assembled a retrospective cohort from seven hospitals, derived a structured 6-hour postoperative feature set, trained three risk models in the development hospitals, and evaluated temporal transport, geographic transport, threshold stability, and silent live operation in external sites. The primary outcome was major postoperative decompensation within 72 hours, defined as unplanned intensive care transfer, septic shock, urgent reoperation for intra-abdominal source control, or death. Models were compared using discrimination, calibration, workload distortion, and threshold regret. Source-optimized boosting achieved the best internal discrimination but showed pronounced external calibration loss and highly unstable alert rates. A domain-regularized model had slightly lower internal performance yet materially better external stability. An ablation analysis demonstrated that observation-process features contributed strongly to internal performance but transported poorly. Silent live operation revealed additional degradation caused by feature latency and backfilled laboratory values. These findings indicate that the primary translational bottleneck was not absence of predictive signal, but failure of score meaning to remain stable outside the environment that generated the training data.