RUN BY JEV
CONNECTING
Connecting the cityLoading official TfL network data…
← Cities

London Underground

Real geography. Simulated operations. AI decisions.

How to play ↗

Simulation connecting…

MORNING RUSH

REAL NETWORK / SIMULATED OPERATIONS

WORLD 02 / METRO / London

London.
Run by Jev.

AI decisions · simulated London

Jev is an AI by TypeSafe, running a simulated London Underground.
Disrupt your copy. See how it responds.

Think you can do better?
Challenge this exact moment.

Moving Crowded Closed
DRAG TO EXPLORE · SCROLL TO ZOOM
NETWORK HEALTH100%
SIMULATED TRIPS0this simulation
CHALLENGE01:30to break the city
YOUR DISRUPTIONS
Three chances. Make them count.

EVIDENCE / LONDON

Can a general AI run the Tube?

CURRENT RESULT / LONDON / SIMULATOR V6 / 22 SEPT 2026

Yes, by a small margin: Jev ran London’s Underground a little better than a purpose-built rulebook.

+0.5 health points on average, and ahead in 48 of 48 matched tests. Each decision takes about 0.3 s and costs $0.0001.

+0.5vs the rulebookWon 48 of 48
+2.2vs doing nothingWon 48 of 48
±0.0vs forecast-onlyNo difference
Average health (0–100, higher is better), zoomed to 84–89 so the gaps are visible

Final health: Jev 86.5 · Forecast-only script 86.6 · Fixed rules 84.8 · No intervention 81.6

How it works. Every 3 seconds Jev gets the city’s state and a 15-second forecast for each option, then chooses. A script that always takes the best forecast scores the same, so the forecasts do most of the work; Jev’s part is reading them, and the situation, in plain language instead of hand-coded rules.

Precisely: Jev was slightly but consistently better than fixed rules: +0.51 points on average (95% interval +0.35 to +0.66). Against fixed rules: 48 wins · 0 ties · 0 losses across 48 matched conditions. Jev was consistently better than no intervention: +2.23 points on average (95% interval +1.99 to +2.46). No reliable difference between Jev and the forecast-only script (always takes the best 15-second forecast): +0.01 points on average (95% interval −0.02 to +0.04). Jev’s briefing includes the same projections; the baseline simply picks the best one.

Caveats & how to read this

The score. Average health is the network health index averaged over the whole 90-second run, the complement of the game’s damage score. It rewards a fast recovery, not just a good final second. Higher is better for the operator. It is a designed congestion index, not an official transit metric.

The ceiling. In 4 of 8 scenarios, Jev and the rulebook both score between 90 and 100: the city has little room to get healthier, so differences there are fractions of a point. Operators separate in the hard scenarios (Station closure, Blackout, Compound, Late pressure), which are highlighted below.

The interval. Differences are paired: each operator faces the same seed, arrivals and disruptions. The 95% interval treats each of the 6 seeds as one observation. An interval that crosses zero means no reliable difference.

Timing. Every operator gets the same decision opportunities; the earlier protocol’s extra final-second action for fixed rules has been removed.

Prompt. metro-choice-4 is live. On these same conditions it changed average health by +0.48 points versus metro-choice-2 (nyc-jev-5 options, no projections) (95% interval +0.31 to +0.64; 48 wins · 0 ties · 0 losses).

Scope. One model run per condition (48 conditions: 6 seeds × 8 scenarios). Trains, riders and disruptions are simulated on real TfL geography; this is not real-world transit performance. “nyc-jev-6” is the version name of the current simulator (simulator v6), used for both cities.

THE DIFFERENCE / PAIRED BY SEED AND SCENARIO

How much does Jev change?#

Jev’s average health minus each alternative Points on a 0–100 scale · dot = mean · bar = 95% interval · right of zero = Jev better
Jev − fixed rules48 wins · 0 ties · 0 losses
+0.51+0.35 to +0.66
Jev − forecast-only script13 wins · 26 ties · 9 losses
+0.01−0.02 to +0.04
Jev − no intervention48 wins · 0 ties · 0 losses
+2.23+1.99 to +2.46
Absolute values · mean across 48 conditions
OperatorAverage healthFinal healthCollapsesReserves left
Jev88.286.563.3
Fixed rules87.784.861.0
Forecast-only script88.286.664.9
No intervention86.081.669.0

EVERY SCENARIO / HARD ONES HIGHLIGHTED

Where does the difference show?#

In 4 of 8 scenarios, Jev and the rulebook both score between 90 and 100: the city has little room to get healthier, so differences there are fractions of a point. Operators separate in the hard scenarios (Station closure, Blackout, Compound, Late pressure), which are highlighted below.

Jev minus each alternative, by scenario Average network health over 90 seconds · 6 seeds per scenario · dot = mean · bar = 95% interval
ScenarioJev − rulesJev − projectionJev − doing nothing
QuietJev 100.0 · rules 99.9 · none 99.6
+0.09
+0.00
+0.42
SignalJev 91.7 · rules 90.8 · none 87.7
+0.86
+0.01
+3.98
Oxford Circus crowdJev 98.5 · rules 98.4 · none 97.7
+0.12
+0.00
+0.75
Stratford crowdJev 99.1 · rules 99.0 · none 98.6
+0.09
−0.01
+0.49
Station closureHARDJev 94.1 · rules 93.4 · none 92.5
+0.73
−0.01
+1.65
BlackoutHARDJev 92.6 · rules 91.9 · none 90.9
+0.61
+0.02
+1.65
CompoundHARDJev 43.5 · rules 42.8 · none 39.4
+0.72
0.00
+4.18
Late pressureHARDJev 86.2 · rules 85.4 · none 81.6
+0.82
+0.04
+4.69
Absolute scores by scenarioAverage health and final health for every operator.
Average network health over 90 seconds (0–100) · final health in brackets
ScenarioJevFixed rulesForecast-only scriptNo intervention
Quiet100.0 (100.0)99.9 (100.0)100.0 (99.8)99.6 (98.5)
Signal91.7 (100.0)90.8 (99.8)91.7 (100.0)87.7 (98.7)
Oxford Circus crowd98.5 (100.0)98.4 (100.0)98.5 (99.8)97.7 (98.5)
Stratford crowd99.1 (100.0)99.0 (99.8)99.1 (100.0)98.6 (98.5)
Station closure · hard94.1 (98.8)93.4 (96.5)94.2 (99.0)92.5 (93.3)
Blackout · hard92.6 (100.0)91.9 (98.7)92.5 (100.0)90.9 (96.5)
Compound · hard43.5 (18.2)42.8 (17.0)43.5 (18.2)39.4 (10.7)
Late pressure · hard86.2 (75.2)85.4 (66.3)86.2 (76.2)81.6 (58.5)

SPEED & COST / CURRENT EVALUATION

Fast enough to run a city?#

785 measured Jev responses from this evaluation. Typical and slow requests both count.

Response time distribution Server round trip, including retries
310 msMedian / p50
645 ms90th percentile
815 ms95th percentile
1,341 ms99th percentile
$0.10Per 1,000 decisions, estimated

785 model requests · $0.079 estimated inference cost · 1,891,781 input tokens · jev-1.13.0. Inference only: hosting and failed-request usage are excluded. A fixed test sample, not a live service SLA.

LIMITS ARE PART OF THE RESULT

Where the city broke.#

Jev: 6 collapses · Fixed rules: 6 collapses · Forecast-only script: 6 collapses · No intervention: 6 collapses in 48 conditions. 0 model request errors and 0 incomplete conditions recorded. A collapse means health stayed critically low; an API error is a separate failure.

Every condition where any operator collapsed · simulated seconds survived
Scenario / seedJevFixed rulesForecast-only scriptNo intervention
Compound · 197723s23s23s23s
Compound · 989623s23s23s23s
Compound · 1781523s23s23s23s
Compound · 2573423s23s23s23s
Compound · 3365323s23s23s23s
Compound · 4157223s23s23s23s

LIVE SHARED CITY

Jev, on the job.#

Connecting to the shared London round…

OFFICIAL TFL / LIVE CONTEXT

What’s happening on the real Tube?#

Arrival predictions and line status refresh while this page is visible. They provide context and do not control simulated trains or demand.

Next trains

Loading TfL arrival predictions…

Refreshes every 30 seconds. Stale predictions are withheld.

Line status

Connecting to TfL…

London’s daily rhythm.

5,283,908
Underground boardingsTypical autumn 2025 Tuesday–Thursday · 05:00–04:59 London time
Boardings throughout a typical weekday 15-minute intervals · historical reference
0500-0515: 2,827 boardings0515-0530: 5,776 boardings0530-0545: 10,479 boardings0545-0600: 17,358 boardings0600-0615: 25,786 boardings0615-0630: 35,456 boardings0630-0645: 46,538 boardings0645-0700: 58,968 boardings0700-0715: 71,806 boardings0715-0730: 85,843 boardings0730-0745: 101,943 boardings0745-0800: 118,233 boardings0800-0815: 130,936 boardings0815-0830: 136,756 boardings0830-0845: 135,934 boardings0845-0900: 127,835 boardings0900-0915: 114,172 boardings0915-0930: 97,333 boardings0930-0945: 83,381 boardings0945-1000: 72,457 boardings1000-1015: 64,320 boardings1015-1030: 57,989 boardings1030-1045: 55,122 boardings1045-1100: 53,759 boardings1100-1115: 52,719 boardings1115-1130: 51,722 boardings1130-1145: 52,072 boardings1145-1200: 53,048 boardings1200-1215: 53,770 boardings1215-1230: 53,631 boardings1230-1245: 54,235 boardings1245-1300: 54,814 boardings1300-1315: 54,864 boardings1315-1330: 54,364 boardings1330-1345: 55,155 boardings1345-1400: 55,989 boardings1400-1415: 56,681 boardings1415-1430: 57,417 boardings1430-1445: 60,005 boardings1445-1500: 63,127 boardings1500-1515: 66,835 boardings1515-1530: 70,989 boardings1530-1545: 78,025 boardings1545-1600: 85,226 boardings1600-1615: 92,659 boardings1615-1630: 99,948 boardings1630-1645: 110,651 boardings1645-1700: 120,420 boardings1700-1715: 129,488 boardings1715-1730: 136,041 boardings1730-1745: 141,194 boardings1745-1800: 139,697 boardings1800-1815: 133,219 boardings1815-1830: 122,954 boardings1830-1845: 111,701 boardings1845-1900: 98,579 boardings1900-1915: 85,879 boardings1915-1930: 74,689 boardings1930-1945: 66,127 boardings1945-2000: 59,104 boardings2000-2015: 53,774 boardings2015-2030: 49,956 boardings2030-2045: 47,748 boardings2045-2100: 45,704 boardings2100-2115: 44,224 boardings2115-2130: 43,291 boardings2130-2145: 43,656 boardings2145-2200: 44,375 boardings2200-2215: 44,901 boardings2215-2230: 44,068 boardings2230-2245: 42,164 boardings2245-2300: 38,410 boardings2300-2315: 33,039 boardings2315-2330: 26,687 boardings2330-2345: 21,066 boardings2345-0000: 16,074 boardings0000-0015: 11,339 boardings0015-0030: 7,323 boardings0030-0045: 4,521 boardings0045-0100: 2,397 boardings0100-0115: 837 boardings0115-0130: 171 boardings0130-0145: 18 boardings0145-0200: 1 boardings0200-0215: 0 boardings0215-0230: 0 boardings0230-0245: 0 boardings0245-0300: 0 boardings0300-0315: 0 boardings0315-0330: 0 boardings0330-0345: 0 boardings0345-0400: 0 boardings0400-0415: 0 boardings0415-0430: 39 boardings0430-0445: 39 boardings0445-0500: 39 boardings
05:0011:0017:0023:0004:59
Northern977,556
District725,014
Victoria708,637
Central683,453
Jubilee672,932
Piccadilly491,358
H&C and Circle411,039
Metropolitan281,906
Bakerloo281,768
Waterloo & City50,246

London Underground line boardings. Hammersmith & City and Circle are combined by TfL. Includes interchange boardings; not unique passengers or gate entries. Historical modelled reference, not live demand or simulator calibration. TfL NUMBAT combines ticketing observations with estimated route choices. The workbook was published 2026-07-01 and retrieved 2026-09-21. Major disruptions and exceptional events are excluded from its typical-day profiles.

SOURCES & LIMITS

Know what’s real.#

OFFICIAL / NETWORK

272 stations. 11 lines.

TfL coordinates and ordered route sequences, retrieved 2026-09-21. Branches are included; reverse duplicate patterns are removed. Lines connect stations, not surveyed tunnels. Elizabeth line, DLR and Overground are not included.

Download the network ↗
MODELLED / SIMULATION

Jev makes the choices.

Demand, train capacities, frequencies, travel times and disruptions are designed simulation parameters. NUMBAT is a reference only. Health is a congestion index; confidence is not a recovery probability.

How the experiment works ↗Download current results ↗
LIVE / TFL

Freshness is visible.

Arrivals use prediction timestamps. Line status shows when TfL was checked, because a separate current-status timestamp is not supplied. Missing feeds are shown as unavailable.

Powered by TfL Open Data ↗

Map backdrop: simplified London borough boundaries and River Thames polygons from the Greater London Authority. Borough boundaries contain Ordnance Survey public sector information licensed under the Open Government Licence v3.0. Registered park boundaries: Crown Copyright 2025 and Ordnance Survey data, released under OGL, via GLA parks data. Major-road context: historical TfL Road Network via GLA. These layers are simplified and do not include every park or street. Some Tube stations lie outside Greater London. Download map geography ↗

Independent experiment. Not affiliated with Transport for London. Each city runs in its own shared world and has its own evaluation; histories and challenge links retain their city identity.

EARLIER STUDIES / KEPT FOR ACCOUNTABILITY

Research history (earlier simulator nyc-jev-5)#

These studies were run before the current evaluation, mostly on the previous simulator, and informed today’s prompt. Their numbers are not the current result above and are not directly comparable with it.

FIRST LONDON EVALUATION / nyc-jev-5 / metro-choice-2 / 2026-09-21

The first London test.

Mean final health (0–100) · 48 conditions · previous simulator
JevFixed rulesNo interventionJev vs rules
96.296.390.27 better · 31 tied · 10 worse

Paired mean difference −0.02 points. Measured before player disruptions were rebalanced; not comparable with the current result.

Download this study ↗

CONTROLLER RESEARCH / NYC-JEV-5

Can better context beat rules?

We tested whether showing Jev approaching trains, blocked reserve approaches and recent passenger transfers improves its choices. The simulator and fixed rules stayed unchanged.

Mean final health / same conditions within each set
Test setCurrent JevMore contextFixed rules
Pilot · 24 conditions96.4697.4296.50
Independent confirmation · 48 conditions96.4296.3896.38

Retain metro-choice-2. The pilot improvement did not reproduce in the independent confirmation set.

One model trial per condition and variant. Final health has limited headroom in these short runs. A richer context also increases input size.

Inspect both experiments ↗

Run your own checks

Inspect decisions or run a fresh comparisonModel outputs, confidence and your own six-condition Jev diagnostic.
JEV / MEASURED IN YOUR CURRENT RUN
0model choices applied

Jev is ready to decide when pressure builds. Start or return to a Jev-operated run to collect your own evidence.

Meet Jev, by TypeSafe AI ↗
Service / selected operator / Rules
Model returned by APIAwaiting first choice
Valid responses / requests0 / 0
Server → Jev latency · p50 / p95 / p99— / — / —
Service / validation errors0
Choice / reported peak mismatches0
Input / output tokens0 / 0
Estimated inference cost$0.000000

Latency includes the server’s TypeSafe round trip and any retries; it is not pure model compute time. Percentiles describe up to 1,200 recent valid responses. Cost uses documented pricing of $0.042 per million input tokens, verified September 21, 2026. Failed-request usage is unavailable; this estimate is not an invoice.

SAME CITY / THREE OPERATORS

Does Jev make the difference?

Six matched scenarios: two seeds, three disruption schedules, 90 simulated seconds each. Jev, rules, and no intervention receive the same arrivals and events. Jev and rules choose from the same valid actions every three ticks.

03
Jev. Rules. No intervention.

Run real model evaluations and inspect survival, health, latency, and token usage. No invented results.

A small simulator diagnostic, not real-world transit validation. A 90-second survival result is capped by the test window. Jev may tie or lose; all completed scenarios are shown. No API calls are made just to display previously recorded results.

DECISION LOG / CURRENT RUN

See the choice. Then the consequence.

0 of 0 matured decisions were followed by higher network health 12 ticks later. This is observational: later decisions and incident recovery also influence the result. It is not causal attribution or confidence calibration.

No model choices in this run yet.

Choose Jev as operator, disrupt a station, then return to inspect its response.

The latest eight records are shown; downloads include up to 500 recent detailed records. Replays store all applied choices. Device-local records and share links are not independently attested benchmark submissions.