hv-drift-rl
A NumPy-only toolkit and benchmark for reinforcement learning under non-stationary drift. Eight mechanisms, six drift shapes, two extra ablations, one meta-agent stub. No PyTorch, no Gym, no dependencies beyond NumPy.
Size: ~20 KB source, no model artifact (stateless). Runtime: ~13 s for the full 9-mechanism ร 6-shape benchmark on a laptop.
The three findings
F1 (refined). Delta encoding is universally safe only when applied to both observation and reward. Delta observation alone is not universally safe: on the corridor task it fails in 3 of 6 drift shapes because the aliased delta state (3 reachable states from 20 positions) cannot carry enough signal to overcome the non-terminal penalty. Adding the reward delta zeroes non-terminal rewards and restores a clean distance-to-goal signal.
F2. Drift shape determines survival, not magnitude. Additive noise
("blur") on the agent's own Q-table is tolerable up to rate 0.05.
Mean-reverting corruption ("decay") collapses at intermediate rates
(0% success at rate 0.01), and partially recovers at higher rates due
to homogenization toward Q.mean().
F3. Drift-as-signal inverts the standard framing. When drift correlates with a hidden regime, reading it beats cancelling it. This is stated in the source chapter but not demonstrated in this prototype โ the corridor task has no hidden regime. See Limitations.
Headline numbers
| mechanism | none | offset | linear | sinusoidal | gaussian | mean_reverting |
|---|---|---|---|---|---|---|
| delta_obs_rew | 100% | 100% | 100% | 100% | 100% | 100% |
| safe_baseline | 100% | 100% | 100% | 100% | 100% | 100% |
| no_correction | 100% | 100% | 12% | 100% | 100% | 100% |
| delta_obs | 0% | 0% | 10% | 90% | 100% | 95% |
| delta_rew_only | 0% | 0% | 0% | 98% | 100% | 100% |
| gym_rms | 100% | 91% | 11% | 99% | 100% | 100% |
| popart | 100% | 100% | 18% | 96% | 100% | 100% |
| reward_clipping | 100% | 100% | 13% | 51% | 100% | 100% |
| anchoring | 100% | 100% | 28% | 100% | 100% | 100% |
30/30 consistency checks pass.
The ablation
The v3 refinement to F1 came from adding two extra mechanisms:
delta_rew_onlyโ delta on reward, absolute observationdelta_obsโ delta on observation, absolute rewarddelta_obs_rewโ delta on both
Results on the shapes where no_correction fails (offset, linear):
| shape | no_correction | delta_obs | delta_rew_only | delta_obs_rew |
|---|---|---|---|---|
| none | 100% | 0% | 0% | 100% |
| offset | 100% | 0% | 0% | 100% |
| linear | 12% | 10% | 0% | 100% |
Neither half alone rescues the failing shape. Only the composite works. This contradicts the source chapter's F1 ("delta obs is universally safe") and produces a stronger claim: the composition is the mechanism, not either half.
The self-drift non-monotonicity
The decay corruption shows a non-monotonic failure curve:
| rate | blur | decay |
|---|---|---|
| 0.000 | 100% | 100% |
| 0.001 | 100% | 100% |
| 0.005 | 100% | 36% |
| 0.010 | 100% | 0% |
| 0.020 | 100% | 53% |
| 0.050 | 98% | 74% |
Decay is worst at intermediate rates (0.01) and partially recovers
at higher rates (0.02โ0.05). The recovery is due to homogenization:
at high decay rates, every Q-value is pulled toward Q.mean(), which
produces a roughly uniform policy. In this corridor, uniform policy
performs reasonably because the reward is dense and the state space is
one-dimensional. The mechanism is real, not a simulator artifact.
Blur never fails. The asymmetry is the finding.
The composition rule
From the source chapter, restated:
# mechanisms = # drift sources. Extra mechanisms don't help.- Never apply EMA to terminal reward.
- Never anchor to a home policy under continued drift.
- If drift correlates with hidden state, read it, not cancel it.
This prototype empirically validates 1โ3. Point 4 requires a hidden-regime task, which this prototype does not include.
How to use
from hv_drift_rl import (
Corridor, make_obs_drift, make_rew_drift,
make_mechanism, train,
)
# Set up a drifting environment
rng = np.random.default_rng(42)
obs_drift = make_obs_drift('linear', rng)
rew_drift = make_rew_drift('linear', rng)
env = Corridor(length=20, obs_drift_fn=obs_drift, rew_drift_fn=rew_drift)
# Pick a mechanism
mech, discretizer = make_mechanism('delta_obs_rew')
# Train
result = train(env, mech, discretizer, n_episodes=500, seed=42)
print(f"Success rate: {result['success_rate']:.1%}")
- Downloads last month
- -