RUN BY JEV
CONNECTING
Connecting the cityLoading official MTA network data…
← Cities

New York subway

Real geography. Simulated operations. AI decisions.

How to play ↗

Simulation connecting…

MORNING RUSH

REAL NETWORK / SIMULATED OPERATIONS

WORLD 01 / METRO / NYC

New York.
Run by Jev.

AI decisions · simulated NYC

Jev is an AI by TypeSafe, running a simulated NYC subway.
Disrupt your copy. See how it responds.

Think you can do better?
Challenge this exact moment.

Moving Crowded Closed
DRAG TO EXPLORE · SCROLL TO ZOOM
NETWORK HEALTH100%
SIMULATED TRIPS0this simulation
CHALLENGE01:30to break the city
YOUR DISRUPTIONS
Three chances. Make them count.

EVIDENCE / NEW YORK

Can a general AI run the subway?

CURRENT RESULT / NEW YORK / SIMULATOR V6 / 22 SEPT 2026

Yes. Jev ran New York’s subway better than a purpose-built rulebook.

+1.3 health points on average, and ahead in 48 of 48 matched tests. Each decision takes about 0.3 s and costs $0.0001.

+1.3vs the rulebookWon 48 of 48
+4.8vs doing nothingWon 48 of 48
±0.0vs forecast-onlyNo difference
Average health (0–100, higher is better), zoomed to 88–95 so the gaps are visible

Final health: Jev 96.6 · Forecast-only script 96.7 · Fixed rules 95.2 · No intervention 89.8

How it works. Every 3 seconds Jev gets the city’s state and a 15-second forecast for each option, then chooses. A script that always takes the best forecast scores the same, so the forecasts do most of the work; Jev’s part is reading them, and the situation, in plain language instead of hand-coded rules.

Precisely: Jev was consistently better than fixed rules: +1.26 points on average (95% interval +1.12 to +1.40). Against fixed rules: 48 wins · 0 ties · 0 losses across 48 matched conditions. Jev was consistently better than no intervention: +4.80 points on average (95% interval +4.46 to +5.14). No reliable difference between Jev and the forecast-only script (always takes the best 15-second forecast): +0.00 points on average (95% interval −0.02 to +0.02). Jev’s briefing includes the same projections; the baseline simply picks the best one.

Caveats & how to read this

The score. Average health is the network health index averaged over the whole 90-second run, the complement of the game’s damage score. It rewards a fast recovery, not just a good final second. Higher is better for the operator. It is a designed congestion index, not an official transit metric.

The ceiling. In 3 of 8 scenarios, Jev and the rulebook both score between 95 and 100: the city has little room to get healthier, so differences there are fractions of a point. Operators separate in the hard scenarios (Bronx crowd, Station closure, Blackout, Compound, Late pressure), which are highlighted below.

The interval. Differences are paired: each operator faces the same seed, arrivals and disruptions. The 95% interval treats each of the 6 seeds as one observation. An interval that crosses zero means no reliable difference.

Timing. Every operator gets the same decision opportunities; the earlier protocol’s extra final-second action for fixed rules has been removed.

Prompt. metro-choice-4 is live. On these same conditions it changed average health by +1.42 points versus metro-choice-3 (nyc-jev-5 options, no projections) (95% interval +1.38 to +1.47; 47 wins · 0 ties · 1 loss).

Scope. One model run per condition (48 conditions: 6 seeds × 8 scenarios). Trains, riders and disruptions are simulated on real MTA geography; this is not real-world transit performance. “nyc-jev-6” is the version name of the current simulator (simulator v6), used for both cities.

THE DIFFERENCE / PAIRED BY SEED AND SCENARIO

How much does Jev change?#

Jev’s average health minus each alternative Points on a 0–100 scale · dot = mean · bar = 95% interval · right of zero = Jev better
Jev − fixed rules48 wins · 0 ties · 0 losses
+1.26+1.12 to +1.40
Jev − forecast-only script13 wins · 18 ties · 17 losses
+0.00−0.02 to +0.02
Jev − no intervention48 wins · 0 ties · 0 losses
+4.80+4.46 to +5.14
Absolute values · mean across 48 conditions
OperatorAverage healthFinal healthCollapsesReserves left
Jev94.496.602.1
Fixed rules93.195.200.0
Forecast-only script94.496.702.8
No intervention89.689.809.0

EVERY SCENARIO / HARD ONES HIGHLIGHTED

Where does the difference show?#

In 3 of 8 scenarios, Jev and the rulebook both score between 95 and 100: the city has little room to get healthier, so differences there are fractions of a point. Operators separate in the hard scenarios (Bronx crowd, Station closure, Blackout, Compound, Late pressure), which are highlighted below.

Jev minus each alternative, by scenario Average network health over 90 seconds · 6 seeds per scenario · dot = mean · bar = 95% interval
ScenarioJev − rulesJev − projectionJev − doing nothing
QuietJev 99.9 · rules 99.8 · none 99.2
+0.15
−0.01
+0.72
SignalJev 96.3 · rules 96.0 · none 93.2
+0.31
−0.01
+3.11
Crowd at Times SquareJev 98.7 · rules 97.7 · none 96.5
+0.94
+0.00
+2.21
Bronx crowdHARDJev 98.3 · rules 98.2 · none 84.3
+0.11
−0.03
+14.00
Station closureHARDJev 95.3 · rules 94.4 · none 93.3
+0.92
+0.03
+2.00
BlackoutHARDJev 93.3 · rules 92.8 · none 91.7
+0.55
+0.02
+1.61
CompoundHARDJev 80.9 · rules 77.5 · none 71.7
+3.43
−0.00
+9.22
Late pressureHARDJev 92.5 · rules 88.9 · none 87.0
+3.66
+0.01
+5.55
Absolute scores by scenarioAverage health and final health for every operator.
Average network health over 90 seconds (0–100) · final health in brackets
ScenarioJevFixed rulesForecast-only scriptNo intervention
Quiet99.9 (99.2)99.8 (98.5)99.9 (99.3)99.2 (96.3)
Signal96.3 (98.8)96.0 (98.7)96.3 (99.3)93.2 (96.2)
Crowd at Times Square98.7 (99.3)97.7 (98.8)98.7 (99.3)96.5 (96.3)
Bronx crowd · hard98.3 (98.8)98.2 (99.0)98.3 (99.0)84.3 (82.0)
Station closure · hard95.3 (99.0)94.4 (97.8)95.3 (99.3)93.3 (94.5)
Blackout · hard93.3 (98.8)92.8 (96.3)93.3 (98.8)91.7 (93.7)
Compound · hard80.9 (94.5)77.5 (91.0)80.9 (94.3)71.7 (83.7)
Late pressure · hard92.5 (84.2)88.9 (81.5)92.5 (84.0)87.0 (76.0)

SPEED & COST / CURRENT EVALUATION

Fast enough to run a city?#

972 measured Jev responses from this evaluation. Typical and slow requests both count.

Response time distribution Server round trip, including retries
327 msMedian / p50
940 ms90th percentile
1,197 ms95th percentile
1,997 ms99th percentile
$0.11Per 1,000 decisions, estimated

972 model requests · $0.10 estimated inference cost · 2,470,472 input tokens · jev-1.13.0. Inference only: hosting and failed-request usage are excluded. A fixed test sample, not a live service SLA.

LIMITS ARE PART OF THE RESULT

Where the city broke.#

Jev: 0 collapses · Fixed rules: 0 collapses · Forecast-only script: 0 collapses · No intervention: 0 collapses in 48 conditions. 0 model request errors and 0 incomplete conditions recorded. A collapse means health stayed critically low; an API error is a separate failure.

No simulated collapses in this sample. That does not establish real-world reliability.

LIVE SHARED CITY

Jev, on the job.#

Connecting to the shared New York round…

OFFICIAL MTA / REAL-WORLD CONTEXT

Outside the simulation.#

Actual arrival predictions, service alerts and historical ridership from New York. These are separate from the simulated trains you see on Watch.

Next trains

Loading MTA arrival predictions…

Auto-refreshes every 30 seconds while this page is visible.

Service alerts

Connecting to the MTA service-alert feed…

New York, by the numbers.

427,004
subway entries in one hourSeptember 9, 2026 · 08:00–09:00 EDT
424 station complexes in the source snapshot
Grand Central-42 St (4,5,6,7,S)14,033
Times Sq-42 St/Port Authority Bus Terminal (1,2,3,7,A,C,E,N,Q,R,W,S)13,987
34 St-Penn Station (1,2,3)8,287
34 St-Penn Station (A,C,E)7,218
Jackson Hts-Roosevelt Av/74 St-Broadway (7,E,F,M,R)5,428
Flushing-Main St (7)5,357

Official MTA hourly estimates, summed across fare classes and payment methods. Historical entries, not live demand, train occupancy or exits; the simulator does not use them to calibrate demand. View the source ↗

DATA PROVENANCE

Know what you’re looking at.#

OFFICIAL / STATIC

The real network

421 mapped station complexes · 28 services. MTA station coordinates and scheduled connections, retrieved 2026-09-21.

Weekday northbound patterns; Staten Island Railway excluded. Lines join stops, not exact track geometry.

MTA source ↗Download network ↗
OFFICIAL / LIVE

Arrivals & service alerts

MTA arrival predictions refresh every 30 seconds here. Service alerts refresh every minute. Feed age and missing coverage are shown.

Real-world context only. These feeds do not drive the simulated trains or demand.

Feed documentation ↗
OFFICIAL / HISTORICAL

Hourly ridership

2026-09-09 · 08:00–09:00 · America/New_York. 427,004 estimated subway entries in the source hour.

Fare entries, not occupancy or unique people. Reference only; not used to calibrate this simulation.

Dataset ↗Download snapshot ↗
DESIGNED / SIMULATED

The experiment

Trains, passengers, crowding, incidents and health are simulated. Jev’s API choices, response times and token usage are recorded.

Health is a congestion index—not an MTA metric. Confidence is not a success rate.

Read the assumptions ↗Download current results ↗

Map backdrop: NYC borough shorelines, NYC Parks properties, and OTI street centerlines via NYC Open Data. Simplified geographic context; not every park or street is shown. Download map geography ↗

EARLIER STUDIES / KEPT FOR ACCOUNTABILITY

Research history (earlier simulator nyc-jev-5)#

These studies were run before the current evaluation, mostly on the previous simulator, and informed today’s prompt. Their numbers are not the current result above and are not directly comparable with it.

PROMPT STUDY / metro-choice-3 VS metro-choice-2 / NYC-JEV-5 / 2026-09-21

Did the new prompt help?

Mean final health (0–100) · 48 matched conditions · previous simulator
Jev · metro-choice-3Jev · metro-choice-2Fixed rulesNo intervention
84.083.483.977.7

The newer prompt gives Jev the immediate crowd effect of each transfer. Against the previous prompt: +0.56 points, 20 better · 19 tied · 9 worse. Versus fixed rules, the seed-cluster 95% interval was −0.64 to +0.89 points, so no reliable difference. Small synthetic sample; the uncertainty interval includes no advantage. Fixed rules receives an extra terminal action under the unchanged benchmark. No claim of reliable superiority.

Original expanded evaluation · nyc-jev-5First prompt (metro-choice-1), every win, tie and failure.

ORIGINAL PROMPT / metro-choice-1 / nyc-jev-5 / jev-1.13.0 / 2026-09-21

192 trials. All outcomes.

12 seeds × 8 scenarios × 2 Jev repeats. Every trial runs for up to 90 simulated seconds with matched rules and no-intervention comparisons. All 192 trials completed; 0 trial failures. Recorded choices reproduced each Jev run’s final health and rider count exactly.

82.2/100Jev · mean final health
83.3Rules · same conditions
77.6No intervention
3,633Actual model requests

Jev improved health over no intervention by 4.5 points. The rules baseline led Jev by 1.2 points. Survival was identical on average (81.6s). This suite does not establish a Jev advantage over rules.

Mean final network health by scenario (0–100)
ScenarioJevRulesNo action
Quiet98.098.796.0
Signal97.198.289.4
Crowd at Times Square97.998.796.0
Bronx crowd97.399.884.3
Station closure95.897.793.9
Blackout95.395.793.1
Compound16.616.813.9
Late pressure59.261.354.5

Against rules: 20 health wins, 63 ties, 109 losses. Request latency: median 258ms, p95 391ms. Estimated model input cost: $0.242 for the suite; infrastructure excluded.

A designed simulator, not a validated transit benchmark. Repeats share conditions. Outcomes depend on the candidate actions, prompt, parameters, and 90-second horizon. Cost uses recorded input tokens and the documented rate at test time. No claim of statistical significance or real-world transit performance.

Download all trial results ↗

Run your own checks

Inspect Jev’s decisionsModel outputs, latency, costs and a six-scenario diagnostic.
JEV / MEASURED IN YOUR CURRENT RUN
0model choices applied

Jev is ready to decide when pressure builds. Start or return to a Jev-operated run to collect your own evidence.

Meet Jev, by TypeSafe AI ↗
Service / selected operator / Rules
Model returned by APIAwaiting first choice
Valid responses / requests0 / 0
Server → Jev latency · p50 / p95 / p99— / — / —
Service / validation errors0
Choice / reported peak mismatches0
Input / output tokens0 / 0
Estimated inference cost$0.000000

Latency includes the server’s TypeSafe round trip and any retries; it is not pure model compute time. Percentiles describe up to 1,200 recent valid responses. Cost uses documented pricing of $0.042 per million input tokens, verified September 21, 2026. Failed-request usage is unavailable; this estimate is not an invoice.

SAME CITY / THREE OPERATORS

Does Jev make the difference?

Six matched scenarios: two seeds, three disruption schedules, 90 simulated seconds each. Jev, rules, and no intervention receive the same arrivals and events. Jev and rules choose from the same valid actions every three ticks.

03
Jev. Rules. No intervention.

Run real model evaluations and inspect survival, health, latency, and token usage. No invented results.

A small simulator diagnostic, not real-world transit validation. A 90-second survival result is capped by the test window. Jev may tie or lose; all completed scenarios are shown. No API calls are made just to display previously recorded results.

DECISION LOG / CURRENT RUN

See the choice. Then the consequence.

0 of 0 matured decisions were followed by higher network health 12 ticks later. This is observational: later decisions and incident recovery also influence the result. It is not causal attribution or confidence calibration.

No model choices in this run yet.

Choose Jev as operator, disrupt a station, then return to inspect its response.

The latest eight records are shown; downloads include up to 500 recent detailed records. Replays store all applied choices. Device-local records and share links are not independently attested benchmark submissions.

Run your own baseline comparisonA local rules-only diagnostic; no Jev requests.

MEASURED HERE / RULE-BASED SIMULATION

A fair starting line.

Twelve paired scenarios. Identical seeds and disruption schedules. Compare the current rule-based operator with no intervention. This measures the simulator’s baselines—not Jev.

12
Same city. Same pressure. Two operators.

Run the comparison locally to see measured survival and final health. No sample results are prefilled.

90 simulated seconds per scenario · 3 seeds × 4 event configurations · generated in your browser · no API latency or model cost implied.

Your loaded session snapshotSeparate from the live city. Inspect or download this session.

YOUR SESSION / SPRINT / RULES

Your loaded snapshot.

100%Network health index
00:00Simulated time survived
0Operator choices
0Simulated completed trips
START OF AVAILABLE METRICS00:00 / HEALTH 0–100

This snapshot updates when you return to Watch. Personal challenges pause while you explore. Network health is a designed congestion index, not an MTA service metric.