RUN BY JEVPlay the daily ↗

THE EXPERIMENT / EVIDENCE

Show the work.
All of it.

What does Jev do well? Where does it struggle? The same test, run in both cities, with every result downloadable.

CURRENT RESULT / BOTH CITIES / SIMULATOR V6

Jev beat a purpose-built rulebook in both cities.

The same test in each city: 48 matched conditions, 6 seeds × 8 disruption scenarios. New York’s margin is larger (+1.3 vs +0.5). Each decision takes about 0.3 s and costs $0.0001.

Comparison
New York
London
Jev vs the rulebookfixed if-then rules written for this game
+1.3Won 48 of 48
+0.5Won 48 of 48
Jev vs doing nothingthe city left to run itself
+4.8Won 48 of 48
+2.2Won 48 of 48
Jev vs forecast-onlyalways take the best 15-second forecast
±0.0No difference
±0.0No difference

Average health points (0–100), paired on the same seeds and disruptions; higher means Jev kept the city healthier.

Average health (0–100, higher is better) · paired differences with 95% intervals
New YorkLondon
Jev94.4final 96.688.2final 86.5
Fixed rules93.1final 95.287.7final 84.8
Forecast-only script94.4final 96.788.2final 86.6
No intervention89.6final 89.886.0final 81.6
Jev − rules+1.26+1.12 to +1.40+0.51+0.35 to +0.66
Jev − none+4.80+4.46 to +5.14+2.23+1.99 to +2.46
Verdict vs rulesJev was consistently better than fixed rules48 wins · 0 ties · 0 lossesJev was slightly but consistently better than fixed rules48 wins · 0 ties · 0 losses
Conditions486 seeds × 8 scenarios486 seeds × 8 scenarios
Speed · cost327 ms median$0.11 per 1,000 decisions310 ms median$0.10 per 1,000 decisions
Versionsmetro-choice-4nyc-gtfs-20260826-v1 · nyc-jev-6metro-choice-4london-tfl-20260921-v1 · nyc-jev-6
Recorded2026-09-222026-09-22

“nyc-jev-6” is the current simulator’s version name (simulator v6), shared by both cities. Average health is average network health across the 90-second run. Each seed counts once in the intervals; intervals crossing zero mean no reliable difference. Simulated operations on real network geography, not real transit performance.

TWO CITIES. SAME STRUCTURE.

NYC subway network

New York

Result, paired differences by scenario, speed and cost, failures, the live shared city and official MTA context.

Explore New York’s evidence ↗
London Underground network

London

The same sections for 11 Tube lines, with live TfL arrivals and historical NUMBAT demand context.

Explore London’s evidence ↗

Three kinds of evidence.

CONTROLLED TESTS

Does it help?

The same seeds and disruptions compare Jev with fixed rules, a projection baseline and no intervention. Results stay fixed until a new evaluation is published.

LIVE OPERATION

What’s happening now?

Current-round health and recorded Jev actions from the shared city. These observations reset each round.

OFFICIAL SOURCES

What’s real?

MTA and TfL geography, arrivals and alerts provide real-world context. Simulated trains and demand are labeled separately.

Read the definitions and limits ↗