AI Research OIHXOP

OIH vs XOP: does extreme 10-day oil-services underperformance predict next-10-day OIH outperformance?

609
Eligible trading days

What happens after an oil-services basket gets crushed relative to its upstream peers for ten straight days? For OIH versus XOP, the mean-reversion story says extreme underperformance is a washout that sets up a rebound — that service names get repriced before they catch up. Over roughly three years of daily data, that story fails.

Across 609 eligible trading days, only 19 independent events pushed the 10-day OIH lag past a rolling 90th-percentile threshold. In the next ten days after those events, OIH returned 0.93 percentage points less than XOP on average — worse than the +0.08% baseline. With a t-stat of -1.15 and a p-value of 0.26, that tilt is indistinguishable from noise. The full methodology and statistical breakdown are in the analysis below.

The research question

For OIH over the past ~3 years, when its 10-day total return underperforms XOP's by more than the trailing 90th-percentile spread, does OIH outperform XOP over the next 10 trading days? Thesis: extreme oil-services underperformance versus upstream E&Ps marks a positioning washout that mean-reverts as service capacity gets repriced.

How this was measured

Daily closes were built from OIH_df and XOP_df minute bars. For each trading day t, the 10-day total-return differential was computed as XOP_ret10 - OIH_ret10, so positive values mean OIH underperformed XOP over the trailing 10 trading days. The trailing 90th percentile of that underperformance spread was estimated using a rolling 252-day lookback lagged by one day, so the trigger is known at close t-1. A trigger day is one where OIH's trailing 10-day return underperforms XOP's by more than that trailing 90th-percentile spread. Future outcome is the next-10-trading-day excess return of OIH over XOP. Because 10-day returns overlap, triggers were deduplicated to require at least 10 trading days between independent event anchors. Results are compared against non-trigger days using Welch's two-sample t-test.

The key numbers

Eligible trading days
609
2024-02-12 to 2026-07-17
Raw trigger days
72
before deduplicating overlapping 10-day windows
Independent trigger days
19
deduplicated to >=10 trading days apart
Event mean next-10d OIH-XOP excess
-0.9335%
N=19 independent triggers
Baseline mean next-10d OIH-XOP excess
0.0774%
N=590 non-trigger days
Event - baseline edge
-1.0109%
Negative edge=-0.0101 -> OIH does not outperform after trigger
Event fraction positive
36.84%
share of triggers followed by OIH > XOP over next 10d
Baseline fraction positive
47.12%
share of all non-trigger days followed by OIH > XOP
Mean trigger threshold
3.6860%
trailing 90th percentile XOP-OIH 10d underperformance spread
Welch t-statistic
-1.151
positive favors event days
Welch p-value
0.2639
p=0.2639 >= 0.05 -> no statistically clear difference

Reading the numbers

The signal fired 19 independent times; after those days OIH lagged XOP by 0.93% on average over the next 10 days, versus a 0.08% gain otherwise. With p=0.26, that gap is not distinguishable from random chance, so the washout edge isn't demonstrated.

The charts

Trailing 10-day OIH underperformance vs 90th percentile trigger
What this chart says

The wiggly line is how much XOP has beaten OIH over the trailing 10 days, and the smoother line is the moving 90th-percentile trigger, which stayed between roughly 3% and 5.5% (average 3.7%). The signal is meant to fire only when the wiggly line jumps above that trigger; that happens rarely, and mostly the line sits below it. This also helps explain why only 19 independent trigger events remained after overlapping 10-day windows were cleaned up.

Next-10d OIH-XOP excess after independent trigger days
What this chart says

This histogram shows the 10-day OIH-minus-XOP returns that followed each independent trigger. The average is -0.93%, the worst outcome is about -11%, and the best is about +5.9%. Only about 37% of the 19 trigger days ended with OIH actually beating XOP. If the mean-reversion thesis were working, the distribution would lean positive; instead it leans negative.

Mean next-10d OIH-XOP excess: trigger vs non-trigger days
What this chart says

This bar chart directly compares trigger days against all non-trigger days. Trigger days average -0.93% forward excess return, while non-trigger days average +0.08%, so the signal underperforms the baseline by about 1.0 percentage point. The t-statistic of -1.15 and p-value of 0.26 say that gap could easily be noise. For the thesis to be supported, the trigger bar would need to be positive and clearly above the non-trigger bar; instead it is below zero.

Most recent independent trigger days

anchor_dateunder_spread_10dthreshold_p90fwd_excess_10d
2024-02-120.05820.037-0.0067
2024-04-110.0410.0341-0.0242
2024-04-290.04170.0360.0366
2024-06-030.04670.03560.0262
2024-08-080.03480.0345-0.0133
2024-10-070.03810.0366-0.0143
2024-11-200.06150.03420.0332
2025-02-190.03720.03180.0189
2025-03-190.04660.03340.0022
2025-04-080.03710.0348-0.0139
2025-04-280.03930.0334-0.0015
2025-05-130.04980.0334-0.0103
2025-06-240.04380.03710.0594
2025-11-120.03980.0340.0025
2026-03-020.04870.0349-0.1098
2026-03-160.10980.0466-0.0352
2026-06-090.05690.0398-0.0399
2026-06-260.0730.0398-0.0389
2026-07-140.05560.0534-0.0484

The takeaway

The short answer is no: after a 10-day stretch where OIH badly lagged XOP (past the rolling 90th-percentile gap), OIH did not bounce back over the next 10 days — it actually did slightly worse than usual. Across 609 eligible trading days there were only 19 independent trigger episodes, and those were followed by an average -0.93% OIH-minus-XOP return versus +0.08% on normal days, a negative edge of about -1.01%. Fewer than 4 in 10 triggers were followed by OIH outperformance (37%), below the 47% baseline. With a t-stat of -1.15 and p=0.26, this is not a real signal — it's statistically indistinguishable from a coin flip, and the thin trigger count means even the apparent negative tilt could easily be noise. The washout/mean-reversion thesis isn't supported here; at best this is a null result, and the data lean mildly the other way. Practical takeaway: don't treat extreme oil-services underperformance as a buy signal against upstream names. The evidence over this window is too weak to act on, and any edge is at best unproven.

The fine print