hv-drift-rl

A NumPy-only toolkit and benchmark for reinforcement learning under non-stationary drift. Eight mechanisms, six drift shapes, two extra ablations, one meta-agent stub. No PyTorch, no Gym, no dependencies beyond NumPy.

Size: ~20 KB source, no model artifact (stateless). Runtime: ~13 s for the full 9-mechanism ร— 6-shape benchmark on a laptop.

The three findings

F1 (refined). Delta encoding is universally safe only when applied to both observation and reward. Delta observation alone is not universally safe: on the corridor task it fails in 3 of 6 drift shapes because the aliased delta state (3 reachable states from 20 positions) cannot carry enough signal to overcome the non-terminal penalty. Adding the reward delta zeroes non-terminal rewards and restores a clean distance-to-goal signal.

F2. Drift shape determines survival, not magnitude. Additive noise ("blur") on the agent's own Q-table is tolerable up to rate 0.05. Mean-reverting corruption ("decay") collapses at intermediate rates (0% success at rate 0.01), and partially recovers at higher rates due to homogenization toward Q.mean().

F3. Drift-as-signal inverts the standard framing. When drift correlates with a hidden regime, reading it beats cancelling it. This is stated in the source chapter but not demonstrated in this prototype โ€” the corridor task has no hidden regime. See Limitations.

Headline numbers

mechanism none offset linear sinusoidal gaussian mean_reverting
delta_obs_rew 100% 100% 100% 100% 100% 100%
safe_baseline 100% 100% 100% 100% 100% 100%
no_correction 100% 100% 12% 100% 100% 100%
delta_obs 0% 0% 10% 90% 100% 95%
delta_rew_only 0% 0% 0% 98% 100% 100%
gym_rms 100% 91% 11% 99% 100% 100%
popart 100% 100% 18% 96% 100% 100%
reward_clipping 100% 100% 13% 51% 100% 100%
anchoring 100% 100% 28% 100% 100% 100%

30/30 consistency checks pass.

The ablation

The v3 refinement to F1 came from adding two extra mechanisms:

  • delta_rew_only โ€” delta on reward, absolute observation
  • delta_obs โ€” delta on observation, absolute reward
  • delta_obs_rew โ€” delta on both

Results on the shapes where no_correction fails (offset, linear):

shape no_correction delta_obs delta_rew_only delta_obs_rew
none 100% 0% 0% 100%
offset 100% 0% 0% 100%
linear 12% 10% 0% 100%

Neither half alone rescues the failing shape. Only the composite works. This contradicts the source chapter's F1 ("delta obs is universally safe") and produces a stronger claim: the composition is the mechanism, not either half.

The self-drift non-monotonicity

The decay corruption shows a non-monotonic failure curve:

rate blur decay
0.000 100% 100%
0.001 100% 100%
0.005 100% 36%
0.010 100% 0%
0.020 100% 53%
0.050 98% 74%

Decay is worst at intermediate rates (0.01) and partially recovers at higher rates (0.02โ€“0.05). The recovery is due to homogenization: at high decay rates, every Q-value is pulled toward Q.mean(), which produces a roughly uniform policy. In this corridor, uniform policy performs reasonably because the reward is dense and the state space is one-dimensional. The mechanism is real, not a simulator artifact.

Blur never fails. The asymmetry is the finding.

The composition rule

From the source chapter, restated:

  1. # mechanisms = # drift sources. Extra mechanisms don't help.
  2. Never apply EMA to terminal reward.
  3. Never anchor to a home policy under continued drift.
  4. If drift correlates with hidden state, read it, not cancel it.

This prototype empirically validates 1โ€“3. Point 4 requires a hidden-regime task, which this prototype does not include.

How to use

from hv_drift_rl import (
    Corridor, make_obs_drift, make_rew_drift,
    make_mechanism, train,
)

# Set up a drifting environment
rng = np.random.default_rng(42)
obs_drift = make_obs_drift('linear', rng)
rew_drift = make_rew_drift('linear', rng)
env = Corridor(length=20, obs_drift_fn=obs_drift, rew_drift_fn=rew_drift)

# Pick a mechanism
mech, discretizer = make_mechanism('delta_obs_rew')

# Train
result = train(env, mech, discretizer, n_episodes=500, seed=42)
print(f"Success rate: {result['success_rate']:.1%}")
Downloads last month
-
Video Preview
loading