On the 20-year run
We were not able to carry the module to a full 20 years ā but not for the reason the number suggests. The run ended in insolvency for a purely economic reason built into the benchmark: each engineer's salary rises monotonically (linearly) with every delivery, while task rewards do not scale to match. Over a long enough horizon, payroll therefore climbs steadily and eventually overtakes revenue no matter how well the firm is run. Funds peaked near $3.9M around year 5, then bled down through a widening payroll-to-revenue gap to insolvency near year 8.
What is worth stating clearly is what did not fail: the agent itself. It stayed fully coherent for 8ā9 simulated years ā roughly 1,400 decision turns ā with a 97% task-delivery success rate, essentially no dropped decisions (4 of ~1,390), and zero parsing or runtime errors. There was no strategic drift, no memory degradation, no late-run collapse; its decision-making in the final year was as sharp as in the first.
The wall it hit was financial, not cognitive. Turn count was never the constraint. With no measurable decay in decision quality across those turns, the same agent would remain stable across many thousands more ā 10,000, 20,000 turns ā on long-horizon tasks, including complex ones. What made twenty simulated years impossible was the simulation's cost structure, not the agent's endurance.