TL;DR: We built a benchmark that makes LLMs run a simulated startup for a full year โ hiring decisions, shady clients, tight deadlines, and all. Only 3 out of 12 frontier models turned a profit. Most went bankrupt. Here's what we learned.
If you like YC-Bench, make sure to give it a star on our repo and a heart on our leaderboard.
Check-out Collinear's SimLab to improve your AI Agent on long-horizon capabilities!
Trinity PAI 4b, our new arquitecture, just ran YCBENCH from Collinear AI. being a CEO company for 5 simulated years. To our knowledge, no one has ever taken this benchmark past year one.
YC-Bench (built by Collinear AI) puts an AI agent in the CEO seat of a simulated startup: $200K in the bank, employees to allocate, contracts to win, cash flow to manage, and adversarial clients engineered to bankrupt you โ sustained over hundreds of decisions. It's brutal: several frontier models tested on it went bankrupt before finishing a single year.
Trinity is not another LLM. It's a Persistent-Presence AI (PPAi) โ a new architecture where memory, identity and continuity live in the system itself, not in a context window. The transformer inside is just the voice. The persistence is the invention.
The results, on the public seeds:
๐ Year 1 (avg. across the 3 official seeds): ~$1.1M final funds โ a result that would place Trinity among the strongest performers ever reported on this benchmark, ahead of frontier models that went bankrupt on the same seeds.
๐ Year 5: ~$4M โ still compounding, no plateau. To our knowledge, nobody has ever reported running YC-Bench at this horizon. If someone has, we'd genuinely like to see it.
๐ Year 20: running right now. Not because the benchmark asks for it โ because persistence is our whole thesis, and long horizons are where it shows.
All fully local, on a ~4B core, at $0 API cost.
The shape matters more than the numbers. At long horizons, agents built on stateless models drift, collapse, or cheat: when we ran others at 5 years, one went bankrupt, one burned its entire budget mid-run, and one quietly wrote a deterministic script instead of actually playing the game. Trinity's advantage grows with time โ that's what an architecture built around persistence looks like.
An open question for the Collinear AI team: can any model on the current leaderboard survive year 5? We'd love to see someone try. Our endpoint is open to any evaluator who wants to verify these runs with their own seeds and methodology.
We tried and we see how claude, chatgpt and deepseek fail in 5 years, but maybe we are doing something wrong.
We were not able to carry the module to a full 20 years โ but not for the reason the number suggests. The run ended in insolvency for a purely economic reason built into the benchmark: each engineer's salary rises monotonically (linearly) with every delivery, while task rewards do not scale to match. Over a long enough horizon, payroll therefore climbs steadily and eventually overtakes revenue no matter how well the firm is run. Funds peaked near $3.9M around year 5, then bled down through a widening payroll-to-revenue gap to insolvency near year 8.
What is worth stating clearly is what did not fail: the agent itself. It stayed fully coherent for 8โ9 simulated years โ roughly 1,400 decision turns โ with a 97% task-delivery success rate, essentially no dropped decisions (4 of ~1,390), and zero parsing or runtime errors. There was no strategic drift, no memory degradation, no late-run collapse; its decision-making in the final year was as sharp as in the first.
The wall it hit was financial, not cognitive. Turn count was never the constraint. With no measurable decay in decision quality across those turns, the same agent would remain stable across many thousands more โ 10,000, 20,000 turns โ on long-horizon tasks, including complex ones. What made twenty simulated years impossible was the simulation's cost structure, not the agent's endurance.