Operating Room Scheduling Efficiency
A simulation and optimization study of 35,772 orthopaedic surgical cases at a multi‑campus academic health system.
Case time Turnover Overrun past block end Schematic of a 540‑minute block day — the unit of analysis throughout this study.
1Overview
This dissertation asks a narrow operational question with a wide institutional answer: why do orthopaedic operating rooms run late, and what would actually fix it? Three years of case‑level Epic OpTime data from a seven‑site academic health system are used to characterize baseline OR performance, build and validate a discrete‑event simulation of the suite, identify the empirical drivers of OR‑day overtime, and test a constraint‑programming scheduling optimizer against both observed practice and a classical greedy heuristic.
Three findings carry the work. First, delay cascades — a case running past the scheduled start of the case behind it — are the dominant structural driver of overtime, worth roughly 103 additional overtime minutes per event at the OR‑day level. Second, a CP‑SAT optimizer eliminated scheduled overtime entirely on a representative week and cut simulated overtime hours per day by 77% relative to observed practice, solving to certified optimality in under ten seconds. Third, and least expected, the institution's block‑governance rules produce a governance paradox: the surgeons whose utilization qualifies them for two‑room access generate the most overtime precisely when that access is withdrawn through room consolidation.
The study is framed within sociotechnical systems theory and behavioral operations management. Scheduling is treated as a joint technical and social system: an optimizer that ignores how block time is governed will not survive contact with the schedule it is meant to improve.
2Setting, data and scope
The empirical foundation is a three‑year census of orthopaedic surgical cases (CY2023–2025) extracted from Epic OpTime at a multi‑campus academic health system. Of 36,744 extracted records, 35,772 completed cases entered the analytic sample after removing 972 cancellations; records with implausible durations or a missing scheduled start were additionally excluded from the analyses that require those fields. Those cases span 13,888 unique OR‑room‑days, 45 attending orthopaedic surgeons, 91 operating rooms and seven sites.
Case volume concentrates in a small number of services. Hip & Knee arthroplasty is the largest at 12,423 cases (34.7% of the sample), followed by arthroscopy (6,163; 17.2%), other orthopaedic (5,134; 14.4%), hand & wrist (4,206; 11.8%) and shoulder & elbow (2,359; 6.6%). Spine is low‑volume and high‑variance: 1,163 cases at a mean of just over five hours each.
| Surgical service | n | Mean | Median | SD | P25 | P75 |
|---|---|---|---|---|---|---|
| Spine | 1,163 | 302.3 | 283.0 | 116.0 | 212.5 | 373.0 |
| Shoulder & Elbow | 2,359 | 182.9 | 176.0 | 71.0 | 137.0 | 216.5 |
| Hip & Knee Arthroplasty | 12,423 | 163.4 | 148.0 | 69.4 | 120.0 | 188.0 |
| Other / Unknown | 161 | 140.5 | 138.0 | 30.8 | 123.0 | 157.0 |
| Soft Tissue / Wound | 782 | 134.2 | 97.5 | 103.4 | 73.0 | 149.0 |
| Fracture / Trauma | 1,466 | 132.6 | 103.0 | 98.6 | 69.0 | 162.0 |
| Foot & Ankle | 1,745 | 120.5 | 107.0 | 61.3 | 78.0 | 146.0 |
| Arthroscopy | 6,163 | 120.4 | 113.0 | 53.2 | 88.0 | 144.0 |
| Other Orthopaedic | 5,134 | 117.7 | 114.0 | 79.9 | 51.0 | 154.0 |
| Hand & Wrist | 4,206 | 86.1 | 67.0 | 63.6 | 38.0 | 118.0 |
| All services | 35,772 | 134.4 | 120.0 | 78.6 | 78.0 | 168.0 |
Every distribution is right‑skewed — the mean exceeds the median in all ten categories — which is why the simulation is driven by fitted log‑normal and gamma families rather than by means.
Site identities are masked throughout this page and in the dissertation itself, in keeping with the de‑identified data package covered by the executed data transfer and use agreement. Sites are designated Hospital A, B and C with their associated ambulatory and pavilion suites; surgeons carry coded identifiers.
3Theoretical framework
Sociotechnical systems
Sociotechnical systems theory, originating with Trist and Bamforth's Tavistock work on the longwall method of coal‑getting (1951), holds that organizational performance is determined by the joint optimization of social and technical subsystems. In perioperative services the technical subsystem is the scheduling algorithm, the EHR configuration, the block‑allocation rules and the physical room layout; the social subsystem is surgeon behavior, nursing team dynamics, administrative oversight and inter‑departmental coordination norms.
The failures documented here — chronic first‑case delays, inflated turnover, systematic duration underestimation — are textbook sociotechnical misalignment. Technically rational scheduling rules are subverted by predictable human responses: block time requested with padding, patients arriving late to holding, coordinators overriding the generated schedule on relationship grounds rather than data. That lens is what justifies treating the algorithmic intervention and the behavioral predictors as complementary levels of analysis rather than as competing explanations.
Behavioral operations management
Behavioral operations management supplies the second half of the frame. Where classical operations research assumes decision‑makers follow optimal scheduling rules, BOM treats deviations from normative models as systematic, institutionally patterned and therefore measurable. First‑case start delay, excess turnover and duration‑estimation error are modeled here not as noise but as structured behavioral deviations whose magnitude is predictable from observable case and room characteristics. Quantifying each deviation class separately converts a generic scheduling problem into a prioritized list of behavioral interventions.
4Research questions and hypotheses
Scope was revised with the co‑chairs in May 2026. The dissertation pursues four research objectives, operationalized as four research questions and seven testable hypotheses.
| Objective | Research question | |
|---|---|---|
| RO1 | Characterize baseline OR performance | RQ1. What were the levels of utilization, first‑case on‑time starts, turnover and overtime, and how did they vary by room, surgeon, service line and day of week? |
| RO2 | Develop and validate a stochastic simulation | RQ2. Could a discrete‑event simulation calibrated to empirically fitted case‑duration distributions reproduce observed utilization and overtime within acceptable validation tolerances? |
| RO3 | Identify empirical predictors of overtime | RQ3. Which case‑ and OR‑day‑level factors predicted overtime, and which block of predictors contributed the largest incremental ΔR²? |
| RO4 | Develop and evaluate a CP‑SAT optimizer | RQ4. Did a CP‑SAT optimizer generate schedules with materially lower overtime than both observed practice and a first‑fit‑decreasing heuristic, evaluated through the calibrated simulation? |
| Statement | Outcome | |
|---|---|---|
| H1 | A simulation calibrated to fitted service‑time distributions reproduces observed KPIs within tolerance | Partial → supported on the elective cohort |
| H2a | CP‑SAT produces lower overtime than observed practice | Supported |
| H2b | CP‑SAT outperforms the first‑fit‑decreasing heuristic | Supported |
| H3 | First‑case delay, turnover and duration deviation predict OR‑day overtime after controls | Supported |
| H4 | Cascade events drive delay propagation among non‑first cases | Supported |
| H5a | Room consolidation is associated with overtime | Supported |
| H5b | The consolidation effect is moderated by surgeon volume | Supported |
The regression component (H3) was pre‑registered on 7 April 2026 — unit of analysis, aggregation rules, model specifications, assumption checks and sensitivity analyses committed to version control before any model was estimated.
5Methods
The research program runs in five phases.
Phase 1 — Characterization
Exploratory analysis of the full census establishes site‑specific baselines for utilization, first‑case on‑time starts (FCOTS), turnover time (TOT) and overtime, by room, surgeon, service line and day of week.
Phase 2 — Distribution fitting
Maximum‑likelihood fitting of normal, log‑normal and gamma families to case duration for each of nine service categories, selected on AIC. Log‑normal wins for five services (hand & wrist, foot & ankle, fracture & trauma, spine, soft tissue); gamma for the remaining four. Turnover time is gamma with shape 4.15 and scale 13.5 min (n = 21,406 consecutive same‑room case pairs), beating both alternatives by more than two AIC units.
Phase 3 — Discrete‑event simulation
A SimPy model of the OR suite driven by the fitted distributions, run over a 365‑day horizon with a fixed seed, validated against observed KPIs (H1) and re‑fitted separately for two sites as a calibration check.
Phase 4 — Constraint‑programming optimization
A CP‑SAT model built on Google OR‑Tools minimizing a weighted sum of overtime minutes (w = 10), unused room‑minutes (w = 1) and service–room affinity violations (w = 50), subject to the institution's overlapping‑surgery constraints. Three policies are compared: P1 observed practice, P2 first‑fit‑decreasing, P3 CP‑SAT, evaluated both statically on a representative week and dynamically through the calibrated simulation.
Phase 5 — Regression
Hierarchical OLS at the OR‑day level with predictor blocks entered sequentially (controls, then case mix, then scheduling deviations), plus a case‑level model of cascade attribution and a consolidation model with a volume interaction. HC3 robust standard errors throughout.
A note on turnover decomposition
Gross turnover time is decomposed into net turnover, idle time and cascade‑attributable time. That decomposition is what makes the cascade mechanism measurable: without it, a room that ran late because the previous case overran is indistinguishable from a room that was simply slow to clean.
6Baseline OR performance
The institution enters the study well below every published benchmark that matters.
Mean first‑case start delay is 15.6 minutes across 13,885 first cases. Mean gross turnover across the 21,406 consecutive same‑room case pairs used for distribution fitting is 56.2 minutes, with a median of 52.0 and a 90th percentile near 94 — a substantial tail of extended room preparation associated with equipment changeovers and deep‑cleaning events. The cohort figures in Table 3 are computed on the cohort stratification sample and differ marginally.
The duration bias is the finding with the longest reach. Mean actual case duration is 134.4 minutes against a mean scheduled duration of 95.3 — and every one of the 45 surgeons in the sample shows positive bias. Individual biases among the ten highest‑volume surgeons range from about +4 to +82 minutes. Universal, one‑directional error of that kind is not idiosyncratic surgeon optimism; it is underestimation embedded in the scheduling templates themselves.
| Location | Mean (%) | Median (%) | OR‑days |
|---|---|---|---|
| Hospital B OR | 70.4 | 75.7 | 6,374 |
| Hospital C Pavilion OR | 68.4 | 73.2 | 501 |
| Hospital A OR | 68.4 | 69.8 | 3,320 |
| Hospital C Ambulatory OR | 59.7 | 64.4 | 91 |
| Hospital B Ambulatory OR | 58.9 | 59.1 | 2,488 |
| Hospital A ASC | 54.3 | 54.4 | 1,101 |
| Hospital A Spine Center | 27.6 | 26.5 | 13 |
| Monday | 70.2 | — | 2,657 |
| Wednesday | 69.1 | — | 3,085 |
| Tuesday | 68.7 | — | 2,526 |
| Friday | 67.6 | — | 2,560 |
| Thursday | 62.7 | — | 2,474 |
| Saturday | 35.0 | — | 286 |
| Sunday | 34.5 | — | 301 |
Panel medians by day of week were not reported in the primary exploratory output. Cancelled cases are excluded.
7Elective and trauma cohorts
The sample contains two operationally distinct case streams that the EHR cannot separate. The Elective/Emergency field is 99.73% "Elective" and is unusable for this purpose, so the trauma cohort was identified at surgeon level from the departmental orthopaedic trauma roster — three surgeons of 45.
| Metric | Elective | Trauma |
|---|---|---|
| Cases | 31,122 | 4,650 |
| Surgeons | 42 | 3 |
| Mean case duration (min) | 137.2 | 192.1 |
| Mean turnover time (min) | 57.0 | 66.0 |
| FCOTS on-time rate (%) | 32.7 | 42.7 |
| Mean first-case delay (min) | 15.1 | 19.3 |
| Overtime rate (%) | 13.3 | 33.0 |
| Cases starting after 17:00 (%) | 1.8 | 14.5 |
| Distinct rooms per weekday | 15.0 | 2.5 |
| Booked minutes per room-day | 361 | 486 |
| Utilization of the 540-min block (%) | 66.9 | 89.9 |
Trauma is over‑represented in overtime by a factor of 2.1
Share of all cases versus share of all overtime cases, CY2023–2025
The capacity contrast is the substantive finding. Trauma rooms book 486 of 540 available block minutes, running at 89.9% utilization against 66.9% in elective rooms. For a scheduled elective service, utilization in the high eighties reads as efficiency. For an arrival‑driven service it reads as the opposite: a block with no reserve absorbs variability by running late rather than by absorbing it internally. Every downstream metric is consistent with that reading.
Trauma volume is also a system‑wide referral concentration. All orthopaedic trauma in the extract was performed at a single site, where it constituted 30.7% of case volume; the other six sites performed effectively none. The capacity question is therefore a system question, not a local one.
Why this matters for everything downstream
Because the two streams behave differently, they are separated before any modeling. The optimizer is scoped to elective cases — trauma arrivals are not schedulable in advance, and assigning them to blocks would misrepresent the decision actually faced. The pooled KPI set also stops being the right comparator for the simulation, which models a single elective regime.
8Simulation validation
The SimPy model is initialized from the fitted duration and turnover distributions and run for 365 simulated days. H1 required MAPE at or below 10% on at least four of five primary KPIs.
| KPI | Obs. | Sim. | MAPE (%) | Obs. | Sim. | MAPE (%) |
|---|---|---|---|---|---|---|
| Mean case duration (min) | 143.5 | 143.8 | 0.2 | 136.3 | 136.7 | 0.3 |
| Mean turnover time (min) | 56.2 | 56.8 | 1.1 | 55.5 | 56.1 | 1.1 |
| Mean cases per room-day | 2.6 | 2.7 | 3.8 | 2.6 | 2.4 | 7.7 |
| FCOTS on-time rate (%) | 34.1 | 38.4 | 12.6 | 32.7 | 31.0 | 5.2 |
| Case-level overtime rate (%) | 15.9 | 10.8 | 32.1 | 13.3 | 6.1 | 54.1 |
| KPIs passing ≤ 10% MAPE | 3 of 5 | 4 of 5 | ||||
Observed values in this table are computed on the simulation validation sample; they differ marginally from the exploratory figures in Tables 1 and 3, which apply their own inclusion rules.
*Observed values in this table are computed on the simulation validation sample; they differ marginally from the exploratory figures in Tables 1 and 3, which apply their own inclusion rules.* The two distributional metrics land within 2% MAPE under both comparators, confirming that the fitted gamma and log‑normal families are correctly parameterized. The failures are architectural, not distributional. In the observed system each surgeon's room runs an independent case list; the pooled‑queue model lets any case land in any room, which spreads load more evenly than reality and under‑produces the concentrated end‑of‑day overruns that arise when an individual list runs long.
Both results are reported. Against the pre‑specified pooled comparator, H1 receives partial support at three of five. Against the elective cohort the model represents, it is supported at four of five. The difference between the two is itself the trauma finding of section 7. Case‑level overtime rate fails under both and remains the open fidelity gap; a per‑room architecture that respects block ownership is the pre‑specified fix.
9Scheduling policy comparison
Three assignment policies are compared. P1 is the historical schedule as executed in Epic OpTime. P2 is first‑fit‑decreasing, the classical greedy heuristic that assigns the longest unscheduled case to the first room with remaining capacity. P3 is the CP‑SAT optimizer.
CP‑SAT cuts overtime; the greedy heuristic makes it worse
365‑day discrete‑event simulation under empirical case demand
Overtime rate (%)
Overtime hours per day
| Metric | P1 Observed | P2 FFD | P3 CP-SAT |
|---|---|---|---|
| Total cases | 155 | 155 | 155 |
| Total overtime (min) | 311 | 2,580 | 0 |
| Total idle time (min) | 10,903 | 2,154 | 12,532 |
| Mean utilization (%) | 64.6 | 90.1 | 61.4 |
Two results deserve emphasis. First, the greedy heuristic is not merely worse than the optimizer — it is worse than doing nothing, producing 8.3 times the overtime of observed practice by over‑compacting the schedule to a nominal 90% utilization. High utilization achieved by removing all slack is a liability in a stochastic system, and P2 is the clean demonstration of why.
Second, P3's idle time increases. That is the intended behavior, not a defect: the objective function penalizes an overtime minute at ten times the cost of an idle one, so the optimizer deliberately buys buffer. The trade is 1,629 additional idle minutes for 311 eliminated overtime minutes on the static week, and a 7.4‑point drop in nominal utilization for a 3.63‑hour daily overtime reduction in the dynamic run.
The solver returned a certified optimal solution in under ten seconds for the benchmark instance, placing it comfortably within the latency budget of any scheduling application. A 3×3 sensitivity grid over the critical‑portion window parameters solved to proven optimality in all nine cells with total overtime invariant, and an independent post‑solve validator found no violation of the critical‑portion, two‑room or short‑case rules in any cell.
10Predictors of overtime
The hierarchical OLS model was estimated on 10,988 OR‑days with overtime minutes as the outcome, entering predictor blocks in the pre‑registered order. The Step 3 block extends the three pre‑registered deviation predictors with cascade count and an all‑cascade‑day indicator, and substitutes net for gross turnover; that extension follows the turnover decomposition adopted after registration and is reported as a documented deviation rather than as the registered specification.
Behavioral scheduling deviations explain most of the variance
Cumulative R² by hierarchical step, n = 10,988 OR‑days
| Predictor | β | Interpretation |
|---|---|---|
| Cascade count (events) | +102.78*** | Each cascade event adds ~103 overtime minutes |
| Mean duration error (min) | +2.11*** | Each minute of booking bias adds 2.11 overtime minutes |
| FCOTS delay (min) | +0.90*** | Each minute of late first start propagates 0.90 minutes |
| Mean net turnover (min) | −0.20*** | Longer planned buffers between cases → less overtime |
| All-cascade day (indicator) | −23.74*** | Separates days where net turnover is undefined |
*** p < .001. The negative turnover coefficient is interpreted as a selection effect: OR‑days with longer average net turnover are days with planned slack, whereas short‑turnover days reflect compressed back‑to‑back scheduling that generates cascades. A parallel logistic model of overtime incidence returns pseudo‑R² = 0.261 and AUC = 0.828.
The cascade mechanism
A cascade is defined as a case whose predecessor in the same room ran past its scheduled start. At the case level, across 21,312 non‑first cases, cascade‑flagged cases started an adjusted 91.5 minutes later than non‑cascade cases (SE = 0.71; R² = 0.538). The unadjusted contrast is starker still: 109.9 minutes against 10.6.
This is the leverage point of the whole study. Because one upstream overrun propagates roughly an hour and a half downstream, a 12‑minute improvement in first‑case readiness that prevents a single cascade returns about 79 net minutes of OR time. No capital investment in the study delivers a comparable return.
11The governance paradox
Block governance at the institution ties two‑room access to a 70% block‑utilization threshold, and consolidates a surgeon's second room when the gap between cases falls below a defined limit. Consolidation is a sensible rule on its face: releasing unused block recovers idle capacity for other work.
Across all 11,304 OR‑days in the consolidation model, that is exactly what it does. Consolidation intensity — minutes of released block time per day — reduces overtime by 1.86 minutes per minute released (p < .001), supporting H5a. The binary flag for whether consolidation occurred at all is not significant, which is itself informative: it is the amount released that matters, not the event.
The interaction term reverses the picture. For surgeons in the high‑volume tier, the consolidation coefficient flips sign and is large and positive: β = +55.98 minutes (p < .001). Restricting the analysis to block‑matched OR‑days and defining the high‑volume tier on cFTE‑adjusted cases per block‑day — a more conservative specification that controls for part‑time clinical appointments — strengthens the interaction to +49.06 minutes against +35.62 under the raw count tier on the same subsample.
The paradox, stated plainly
Eligibility for a second room is earned by demonstrating high block utilization. Demonstrating high block utilization requires the room capacity that eligibility withholds. The surgeons whose caseload most justifies dual‑room access are the ones whose overtime rises most sharply when consolidation takes that access away — and the governance metric that triggers consolidation is computed on the backward‑looking average that consolidation itself depresses.
Stated generally, this is a structural resource‑allocation failure in which eligibility criteria withhold the resource required to demonstrate eligibility. The dissertation names and operationalizes it as the governance paradox, and it is the study's most novel substantive contribution. The specific thresholds are institutional, but the shape of the rule — earn access by proving you already use capacity efficiently — is common enough in block governance that the pattern is unlikely to be unique to this system.
12Duration prediction (supplementary)
Machine‑learning case‑duration models were developed as a supplementary analysis. They do not test a stated hypothesis, and predictions exist only for CY2025, the held‑out test year.
| Model | MAPE (%) | MAE (min) | RMSE (min) | R² |
|---|---|---|---|---|
| Historical mean | 44.4 | 44.2 | 65.7 | 0.441 |
| Random forest | 32.5 | 33.0 | 51.3 | 0.660 |
| XGBoost | 21.2 | 23.2 | 40.4 | 0.789 |
Halving the prediction error relative to the historical mean matters because duration error is the input to the cascade mechanism. Replacing the scheduled‑duration input to the optimizer with predicted durations is a pre‑specified robustness extension that the current study did not complete.
13Practical recommendations
1. Deploy CP‑SAT scheduling for new‑week release
Integrate the optimizer as a decision‑support layer rather than a mandatory assignment engine. A warm‑started solver seeded with the prior day's partial solution builds on existing block history instead of discarding it. A 60‑second nightly re‑optimization against the updated case queue would recover an estimated 3.63 overtime hours per day for the ten‑room configuration modeled here.
No institutional cost data were available to value that reduction. Using a published benchmark for the wage‑and‑benefit share of OR cost — roughly $13–14 per OR minute, on the order of $800 per OR hour — 3.63 hours per day across 250 operating days implies annual savings on the order of $0.7–0.8 million for that ten‑room configuration. The figure uses a regular‑time wage rate and therefore excludes any overtime premium, and it applies only to the rooms represented in the simulation, not to the system's full 91.
2. Institute a first‑case readiness protocol
Pre‑anesthesia in holding before 07:15, transport confirmed by 07:25, surgeon scrubbed by 07:30. That structure could recover 10–12 minutes of the observed 15.6‑minute mean start delay. Given the 91.5‑minute cascade leverage ratio, a 12‑minute improvement that prevents one cascade per day returns roughly 79 net minutes — the highest‑return process investment available inside the existing scheduling infrastructure.
3. Revise consolidation policy for high‑volume surgeons
Three changes follow directly from H5b. Normalize the utilization denominator on cFTE so part‑time clinical faculty are judged on per‑block performance rather than total volume. Stratify by tenure, granting provisional dual‑room access during the first 18 months of appointment pending 12 months of utilization data. And replace the backward‑looking utilization average with prospective, simulation‑based eligibility projection — evaluating each OR day on its specific case mix rather than on the surgeon's historical record. Together these eliminate the self‑referential access barrier at the heart of the paradox.
4. Run a quarterly duration‑bias audit
Epic OpTime allows procedure‑specific scheduled‑duration defaults to be updated by administrators. Comparing mean actual to mean scheduled duration by CPT code within each service line on a quarterly cycle gives the operational mechanism. Spine (+82.2 min bias) and hip & knee (+38.5 min) are the highest‑priority targets; hand & wrist is already nearly unbiased. Correcting half the systemwide bias would reduce expected cascade‑attributable overtime by roughly 40 minutes per OR‑day.
5. Reassess orthopaedic trauma capacity
The trauma service is not scheduling badly; it is absorbing more work than its allocated room‑hours can hold. At 89.9% block utilization on 2.5 rooms per weekday, it has no reserve with which to absorb arrival variability, and it discharges that variability as overtime and after‑17:00 starts. The appropriate lever is an allocation review — additional dedicated block, or a formal reserved‑capacity arrangement — not a scheduling‑discipline intervention. The efficiency headroom this study identifies sits in the elective rooms at 66.9%, which is where the optimizer is scoped.
14Contributions
A unified methodological pipeline. Simulation, optimization and empirical regression are usually pursued as separate traditions. Here, empirical distribution fitting supplies the stochastic inputs to the simulation; the simulation supplies a validated computational laboratory for policy experiments; and the regression identifies which behavioral deviations to target. Each stage feeds the next rather than standing alone.
Scale and institutional context. Most empirical studies in this literature use fewer than 2,000 case observations, often over less than twelve months, and typically from community hospitals or single‑specialty centers. This study uses a three‑year census from a large U.S. academic medical center with a multi‑site portfolio, resident and fellow participation, and integrated inpatient and ambulatory streams.
The governance paradox as a construct. The study formally states, operationalizes and empirically tests a structural failure mode in resource‑allocation governance, and supplies the quantitative mechanism behind it.
Implementation orientation. Findings are expressed in the metrics the institution already manages by — first‑case on‑time starts, turnover minutes, overtime minutes — and framed as a quality‑improvement initiative within the investigator's operational purview, addressing the persistent criticism that operations management prescriptions are analytically sophisticated and rarely implemented.
15Limitations
Perfect‑foresight durations in the optimizer. The most consequential caveat. P1 derives from historical Epic schedules, which carry the +39.1‑minute booking bias. The optimizer was solved against realized durations — each case's true length, information no scheduler holds at assignment time. This flatters P3. The comparison stays internally valid because all three policies are evaluated on the same realized durations, so the ranking of assignment rules is unaffected, but the magnitude of P3's advantage is an upper bound. Re‑solving with scheduled durations, and again with ML‑predicted durations, would bound it properly; both are pre‑specified extensions not yet completed.
Pooled‑queue simulation architecture. The model treats any case as assignable to any room, which inflates simulated utilization and suppresses FCOTS variance. This limits the simulation's use as a site‑level forecasting tool, though not as a policy‑comparison instrument. A per‑room architecture requires the master block schedule.
Single health system. Generalizability to other specialties, institutions and governance frameworks is untested. The governance paradox in particular depends on this institution's specific two‑room guidelines — the 70% threshold, the 90‑minute gap rule, the escalating penalty structure — which may differ substantially elsewhere.
Critical portion inferred, not observed. The overlapping‑surgery constraints turn on the critical portion of each case, which Epic OpTime does not timestamp. It was inferred from incision and close times as the middle half of operative time. A 3×3 sensitivity grid shows the policy conclusion is insensitive to the window within the range tested, but the constraint remains an inference from surrogate timestamps. Direct capture of critical‑portion start and end times is the single most useful data addition for future work on overlapping surgery.
De‑identification and surgeon‑level power. Surgeon identifiers were coded before analysis, which preserves the exempt determination but precludes individual‑level follow‑up. With 45 surgeons, power for between‑surgeon analysis is limited; the formal bimodality coefficient for annual case volume (0.477) sits below the conventional 0.556 threshold, so the volume stratification underpinning H5b rests on the interaction term rather than on a formally bimodal distribution.
Collinearity. Elective percentage and mean ASA are near‑constant in an orthopaedic sample and are collinear with other scheduling‑density covariates. Both were retained as controls; the pre‑analysis plan set an exclusion rule at a variance inflation factor above 10, so their retention is a documented deviation rather than a pre‑specified choice. The scheduling‑deviation predictors that carry the hypotheses sit well below the conventional threshold.
16Future research
- Per‑room simulation architecture respecting block ownership, surgeon‑to‑room assignment and equipment availability — the direct fix for the remaining fidelity gap, and the prerequisite for site‑level utilization forecasting.
- ML‑driven duration inputs to the optimizer, replacing scheduled durations with predictions conditioned on CPT code, surgeon, patient factors and case order within the block.
- Real‑time re‑optimization triggered when a case overruns by a defined threshold, addressing cascade propagation at the intraday operational level rather than at the weekly scheduling level.
- Multi‑site and multi‑specialty replication, particularly in specialties with greater case‑complexity variance, to test whether the cascade coefficient and the consolidation moderation are orthopaedic features or general ones.
- Natural experiment on governance enforcement. The escalating penalty ladder for two‑room violations is a pre‑coded policy shock; a difference‑in‑differences design around an enforcement event would move H5a/H5b from association to causal identification.
- EHR integration, embedding simulation‑based consolidation recommendations and two‑room eligibility projections directly in the scheduling workflow.
17Reproducibility and data handling
Analysis is script‑based and version‑controlled. The stack is Python — pandas, NumPy, statsmodels, scipy, scikit‑learn — with SimPy for discrete‑event simulation and Google OR‑Tools CP‑SAT for optimization; random seeds are fixed. The H3 pre‑analysis plan was committed before estimation and specifies aggregation rules, model forms, assumption checks and four sensitivity analyses, together with a deviations protocol requiring any departure to be documented and flagged as exploratory.
Every optimizer run is checked with an independent post‑solve validator. A solver returning OPTIMAL means only that the model it was given was solved — not that the model encoded the intended constraints.
Data are de‑identified for quality‑improvement and dissertation purposes under an executed reciprocal data transfer and use agreement. Mapping keys are retained at the source institution and are not part of the transferred package. Case‑level data are not publicly available.
18Selected references
The dissertation's reference list runs to 104 entries. The following are the works most load‑bearing for the arguments summarized above.
Scheduling models and reviews
- Cardoen, B., Demeulemeester, E., & Beliën, J. (2010). Operating room planning and scheduling: A literature review. European Journal of Operational Research, 201(3).
- Samudra, M., Van Riet, C., Demeulemeester, E., et al. (2016). Scheduling operating rooms: Achievements, challenges and pitfalls. Journal of Scheduling, 19.
- Zhu, S., Fan, W., Yang, S., et al. (2019). Operating room planning and surgical case scheduling: A review of literature. Journal of Combinatorial Optimization, 37.
- Denton, B. T., Miller, A. J., Balasubramanian, H. J., & Huschka, T. R. (2010). Optimal allocation of surgery blocks to operating rooms under uncertainty. Operations Research, 58(4).
- Dexter, F., & Traub, R. D. (2002). How to schedule elective surgical cases into specific operating rooms to maximize the efficiency of use of operating room time. Anesthesia & Analgesia, 94(4).
- Wachtel, R. E., & Dexter, F. (2008). Tactical increases in operating room block time for capacity planning should not be based on utilization. Anesthesia & Analgesia, 106(1).
Simulation and stochastic modeling
- Strum, D. P., May, J. H., & Vargas, L. G. (2000). Modeling the uncertainty of surgical procedure times: Comparison of log-normal and normal models. Anesthesiology, 92(4).
- May, J. H., Strum, D. P., & Vargas, L. G. (2000). Fitting the lognormal distribution to surgical procedure times. Decision Sciences, 31(1).
- Hassanzadeh, H., Boyle, J., Khanna, S., et al. (2023). A discrete event simulation for improving operating theatre efficiency. International Journal of Health Planning and Management, 38(2).
- Law, A. M. (2007). Simulation Modeling and Analysis (4th ed.). McGraw-Hill.
Optimization and constraint programming
- Perron, L., & Furnon, V. (2019). OR-Tools. Google Optimization Tools.
- Laborie, P., Rogerie, J., Shaw, P., & Vilím, P. (2018). IBM ILOG CP Optimizer for scheduling. INFORMS Journal on Computing, 30(1).
- Addis, B., Carello, G., Grosso, A., & Tanfani, E. (2016). Operating room scheduling and rescheduling: A rolling horizon approach. Flexible Services and Manufacturing Journal, 28.
- Fairley, M., Scheinker, D., & Brandeau, M. L. (2019). Improving the efficiency of the operating room environment with an optimization and machine learning model. Health Care Management Science, 22.
Duration prediction
- Zaribafzadeh, H., Webster, W. L., Vail, C. J., et al. (2023). Development, deployment, and implementation of a machine learning surgical case length prediction model. Annals of Surgery, 278(6).
- Bozkurt, S., Zeng, C., Koepcke, L., et al. (2024). Predicting surgical case duration using machine learning: A multi-institutional study. Journal of the American College of Surgeons.
- Lex, J. R., Abbas, A., Mosseri, J., et al. (2025). Using machine learning to predict-then-optimize elective orthopedic surgery scheduling. JMIR Medical Informatics.
First-case starts, turnover and cost
- Knox, C., Harper, J., McMillan, L., et al. (2024). Increasing first case on-time starts in the operating room using an electronic readiness dashboard. Perioperative Care and Operating Room Management.
- MacMillan, L., Madura, G. M., Elliot, M., et al. (2025). What affects operating room turnover time? A systematic review and mapping of the evidence. Surgery.
- Childers, C. P., & Maggard-Gibbons, M. (2018). Understanding costs of care in the operating room. JAMA Surgery, 153(4).
- Dexter, F., & Epstein, R. H. (2009). Typical savings from each minute reduction in tardy first case of the day starts. Anesthesia & Analgesia, 108(4).
- Rothstein, D. H., & Raval, M. V. (2018). Operating room efficiency. Seminars in Pediatric Surgery, 27(2).
Theory
- Trist, E. L., & Bamforth, K. W. (1951). Some social and psychological consequences of the longwall method of coal-getting. Human Relations, 4(1).
- Cherns, A. (1976). The principles of sociotechnical design. Human Relations, 29(8).
- Carayon, P., Bass, E., Bellandi, T., et al. (2011). Socio-technical systems analysis in health care: A research agenda. IIE Transactions on Healthcare Systems Engineering, 1(3).
- Gino, F., & Pisano, G. (2008). Toward a theory of behavioral operations. Manufacturing & Service Operations Management, 10(4).
- Bendoly, E., Donohue, K., & Schultz, K. L. (2006). Behavior in operations management: Assessing recent findings and revisiting old assumptions. Journal of Operations Management, 24(6).
- KC, D. S., & Terwiesch, C. (2009). Impact of workload on service time and patient safety: An econometric analysis of hospital operations. Management Science, 55(9).
- Fügener, A., Schiffels, S., & Kolisch, R. (2017). Overutilization and underutilization of operating rooms: Insights from behavioral health care operations management. Health Care Management Science, 20.
- Hevner, A. R., March, S. T., Park, J., & Ram, S. (2004). Design science in information systems research. MIS Quarterly, 28(1).
19Citation and contact
Duncan, A. (in progress). Operating room scheduling efficiency: A simulation and optimization approach. Doctoral dissertation, Florida Atlantic University.
I am glad to discuss the methods, the CP‑SAT formulation, or the governance findings with perioperative leaders and researchers working on similar problems. Case‑level data cannot be shared, but the modeling approach is designed to transfer: any academic medical center running Epic OpTime or a comparable EHR with case‑level timestamps can apply the distribution fitting, simulation architecture, optimizer formulation and regression specification with minimal adaptation.