Does it help?
The same seeds and disruptions compare Jev with fixed rules, a projection baseline and no intervention. Results stay fixed until a new evaluation is published.
THE EXPERIMENT / EVIDENCE
What does Jev do well? Where does it struggle? The same test, run in both cities, with every result downloadable.
CURRENT RESULT / BOTH CITIES / SIMULATOR V6
The same test in each city: 48 matched conditions, 6 seeds × 8 disruption scenarios. New York’s margin is larger (+1.3 vs +0.5). Each decision takes about 0.3 s and costs $0.0001.
Average health points (0–100), paired on the same seeds and disruptions; higher means Jev kept the city healthier.
| New York | London | |
|---|---|---|
| Jev | 94.4final 96.6 | 88.2final 86.5 |
| Fixed rules | 93.1final 95.2 | 87.7final 84.8 |
| Forecast-only script | 94.4final 96.7 | 88.2final 86.6 |
| No intervention | 89.6final 89.8 | 86.0final 81.6 |
| Jev − rules | +1.26+1.12 to +1.40 | +0.51+0.35 to +0.66 |
| Jev − none | +4.80+4.46 to +5.14 | +2.23+1.99 to +2.46 |
| Verdict vs rules | Jev was consistently better than fixed rules48 wins · 0 ties · 0 losses | Jev was slightly but consistently better than fixed rules48 wins · 0 ties · 0 losses |
| Conditions | 486 seeds × 8 scenarios | 486 seeds × 8 scenarios |
| Speed · cost | 327 ms median$0.11 per 1,000 decisions | 310 ms median$0.10 per 1,000 decisions |
| Versions | metro-choice-4nyc-gtfs-20260826-v1 · nyc-jev-6 | metro-choice-4london-tfl-20260921-v1 · nyc-jev-6 |
| Recorded | 2026-09-22 | 2026-09-22 |
“nyc-jev-6” is the current simulator’s version name (simulator v6), shared by both cities. Average health is average network health across the 90-second run. Each seed counts once in the intervals; intervals crossing zero mean no reliable difference. Simulated operations on real network geography, not real transit performance.
TWO CITIES. SAME STRUCTURE.
Result, paired differences by scenario, speed and cost, failures, the live shared city and official MTA context.
Explore New York’s evidence ↗The same sections for 11 Tube lines, with live TfL arrivals and historical NUMBAT demand context.
Explore London’s evidence ↗The same seeds and disruptions compare Jev with fixed rules, a projection baseline and no intervention. Results stay fixed until a new evaluation is published.
Current-round health and recorded Jev actions from the shared city. These observations reset each round.
MTA and TfL geography, arrivals and alerts provide real-world context. Simulated trains and demand are labeled separately.