Jev is an AI by TypeSafe, running a simulated London Underground. Disrupt your copy. See how it responds.
↗
Think you can do better? Challenge this exact moment.
Moving Crowded Closed
DRAG TO EXPLORE · SCROLL TO ZOOM
NETWORK HEALTH100%
SIMULATED TRIPS0this simulation
CHALLENGE01:30to break the city
YOUR DISRUPTIONS
Three chances. Make them count.
00:00
01:30
EVIDENCE / LONDON
Can a general AI run the Tube?
CURRENT RESULT / LONDON / SIMULATOR V6 / 22 SEPT 2026
Yes, by a small margin: Jev ran London’s Underground a little better than a purpose-built rulebook.
+0.5 health points on average, and ahead in 48 of 48 matched tests. Each decision takes about 0.3 s and costs $0.0001.
+0.5vs the rulebookWon 48 of 48
+2.2vs doing nothingWon 48 of 48
±0.0vs forecast-onlyNo difference
Average health (0–100, higher is better), zoomed to 84–89 so the gaps are visible
Jev
88.2
Forecast-only script
88.2
Fixed rules
87.7
No intervention
86.0
848586878889
Final health: Jev 86.5 · Forecast-only script 86.6 · Fixed rules 84.8 · No intervention 81.6
How it works. Every 3 seconds Jev gets the city’s state and a 15-second forecast for each option, then chooses. A script that always takes the best forecast scores the same, so the forecasts do most of the work; Jev’s part is reading them, and the situation, in plain language instead of hand-coded rules.
Precisely: Jev was slightly but consistently better than fixed rules: +0.51 points on average (95% interval +0.35 to +0.66). Against fixed rules: 48 wins · 0 ties · 0 losses across 48 matched conditions. Jev was consistently better than no intervention: +2.23 points on average (95% interval +1.99 to +2.46). No reliable difference between Jev and the forecast-only script (always takes the best 15-second forecast): +0.01 points on average (95% interval −0.02 to +0.04). Jev’s briefing includes the same projections; the baseline simply picks the best one.
The score. Average health is the network health index averaged over the whole 90-second run, the complement of the game’s damage score. It rewards a fast recovery, not just a good final second. Higher is better for the operator. It is a designed congestion index, not an official transit metric.
The ceiling. In 4 of 8 scenarios, Jev and the rulebook both score between 90 and 100: the city has little room to get healthier, so differences there are fractions of a point. Operators separate in the hard scenarios (Station closure, Blackout, Compound, Late pressure), which are highlighted below.
The interval. Differences are paired: each operator faces the same seed, arrivals and disruptions. The 95% interval treats each of the 6 seeds as one observation. An interval that crosses zero means no reliable difference.
Timing. Every operator gets the same decision opportunities; the earlier protocol’s extra final-second action for fixed rules has been removed.
Prompt. metro-choice-4 is live. On these same conditions it changed average health by +0.48 points versus metro-choice-2 (nyc-jev-5 options, no projections) (95% interval +0.31 to +0.64; 48 wins · 0 ties · 0 losses).
Scope. One model run per condition (48 conditions: 6 seeds × 8 scenarios). Trains, riders and disruptions are simulated on real TfL geography; this is not real-world transit performance. “nyc-jev-6” is the version name of the current simulator (simulator v6), used for both cities.
In 4 of 8 scenarios, Jev and the rulebook both score between 90 and 100: the city has little room to get healthier, so differences there are fractions of a point. Operators separate in the hard scenarios (Station closure, Blackout, Compound, Late pressure), which are highlighted below.
Jev minus each alternative, by scenario Average network health over 90 seconds · 6 seeds per scenario · dot = mean · bar = 95% interval
785 measured Jev responses from this evaluation. Typical and slow requests both count.
Response time distribution Server round trip, including retries
0
0–100ms
0
100–200ms
365
200–300ms
258
300–500ms
137
500–1kms
25
1k+ms
310 msMedian / p50
645 ms90th percentile
815 ms95th percentile
1,341 ms99th percentile
$0.10Per 1,000 decisions, estimated
785 model requests · $0.079 estimated inference cost · 1,891,781 input tokens · jev-1.13.0. Inference only: hosting and failed-request usage are excluded. A fixed test sample, not a live service SLA.
Jev: 6 collapses · Fixed rules: 6 collapses · Forecast-only script: 6 collapses · No intervention: 6 collapses in 48 conditions. 0 model request errors and 0 incomplete conditions recorded. A collapse means health stayed critically low; an API error is a separate failure.
Every condition where any operator collapsed · simulated seconds survived
Arrival predictions and line status refresh while this page is visible. They provide context and do not control simulated trains or demand.
Next trains
Loading TfL arrival predictions…
Refreshes every 30 seconds. Stale predictions are withheld.
Line status
Connecting to TfL…
London’s daily rhythm.
5,283,908
Underground boardingsTypical autumn 2025 Tuesday–Thursday · 05:00–04:59 London time
Boardings throughout a typical weekday 15-minute intervals · historical reference
05:0011:0017:0023:0004:59
Northern977,556
District725,014
Victoria708,637
Central683,453
Jubilee672,932
Piccadilly491,358
H&C and Circle411,039
Metropolitan281,906
Bakerloo281,768
Waterloo & City50,246
London Underground line boardings. Hammersmith & City and Circle are combined by TfL. Includes interchange boardings; not unique passengers or gate entries. Historical modelled reference, not live demand or simulator calibration. TfL NUMBAT combines ticketing observations with estimated route choices. The workbook was published 2026-07-01 and retrieved 2026-09-21. Major disruptions and exceptional events are excluded from its typical-day profiles.
TfL coordinates and ordered route sequences, retrieved 2026-09-21. Branches are included; reverse duplicate patterns are removed. Lines connect stations, not surveyed tunnels. Elizabeth line, DLR and Overground are not included.
Demand, train capacities, frequencies, travel times and disruptions are designed simulation parameters. NUMBAT is a reference only. Health is a congestion index; confidence is not a recovery probability.
Arrivals use prediction timestamps. Line status shows when TfL was checked, because a separate current-status timestamp is not supplied. Missing feeds are shown as unavailable.
Map backdrop: simplified London borough boundaries and River Thames polygons from the Greater London Authority. Borough boundaries contain Ordnance Survey public sector information licensed under the Open Government Licence v3.0. Registered park boundaries: Crown Copyright 2025 and Ordnance Survey data, released under OGL, via GLA parks data. Major-road context: historical TfL Road Network via GLA. These layers are simplified and do not include every park or street. Some Tube stations lie outside Greater London. Download map geography ↗
Independent experiment. Not affiliated with Transport for London. Each city runs in its own shared world and has its own evaluation; histories and challenge links retain their city identity.
These studies were run before the current evaluation, mostly on the previous simulator, and informed today’s prompt. Their numbers are not the current result above and are not directly comparable with it.
FIRST LONDON EVALUATION / nyc-jev-5 / metro-choice-2 / 2026-09-21
The first London test.
Mean final health (0–100) · 48 conditions · previous simulator
Jev
Fixed rules
No intervention
Jev vs rules
96.2
96.3
90.2
7 better · 31 tied · 10 worse
Paired mean difference −0.02 points. Measured before player disruptions were rebalanced; not comparable with the current result.
We tested whether showing Jev approaching trains, blocked reserve approaches and recent passenger transfers improves its choices. The simulator and fixed rules stayed unchanged.
Mean final health / same conditions within each set
Test set
Current Jev
More context
Fixed rules
Pilot · 24 conditions
96.46
97.42
96.50
Independent confirmation · 48 conditions
96.42
96.38
96.38
Retain metro-choice-2. The pilot improvement did not reproduce in the independent confirmation set.
One model trial per condition and variant. Final health has limited headroom in these short runs. A richer context also increases input size.
Latency includes the server’s TypeSafe round trip and any retries; it is not pure model compute time. Percentiles describe up to 1,200 recent valid responses. Cost uses documented pricing of $0.042 per million input tokens, verified September 21, 2026. Failed-request usage is unavailable; this estimate is not an invoice.
SAME CITY / THREE OPERATORS
Does Jev make the difference?
Six matched scenarios: two seeds, three disruption schedules, 90 simulated seconds each. Jev, rules, and no intervention receive the same arrivals and events. Jev and rules choose from the same valid actions every three ticks.
03
Jev. Rules. No intervention.
Run real model evaluations and inspect survival, health, latency, and token usage. No invented results.
A small simulator diagnostic, not real-world transit validation. A 90-second survival result is capped by the test window. Jev may tie or lose; all completed scenarios are shown. No API calls are made just to display previously recorded results.
DECISION LOG / CURRENT RUN
See the choice. Then the consequence.
0 of 0 matured decisions were followed by higher network health 12 ticks later. This is observational: later decisions and incident recovery also influence the result. It is not causal attribution or confidence calibration.
No model choices in this run yet.
Choose Jev as operator, disrupt a station, then return to inspect its response.
The latest eight records are shown; downloads include up to 500 recent detailed records. Replays store all applied choices. Device-local records and share links are not independently attested benchmark submissions.