Jev is an AI by TypeSafe, running a simulated NYC subway. Disrupt your copy. See how it responds.
↗
Think you can do better? Challenge this exact moment.
Moving Crowded Closed
DRAG TO EXPLORE · SCROLL TO ZOOM
NETWORK HEALTH100%
SIMULATED TRIPS0this simulation
CHALLENGE01:30to break the city
YOUR DISRUPTIONS
Three chances. Make them count.
00:00
01:30
EVIDENCE / NEW YORK
Can a general AI run the subway?
CURRENT RESULT / NEW YORK / SIMULATOR V6 / 22 SEPT 2026
Yes. Jev ran New York’s subway better than a purpose-built rulebook.
+1.3 health points on average, and ahead in 48 of 48 matched tests. Each decision takes about 0.3 s and costs $0.0001.
+1.3vs the rulebookWon 48 of 48
+4.8vs doing nothingWon 48 of 48
±0.0vs forecast-onlyNo difference
Average health (0–100, higher is better), zoomed to 88–95 so the gaps are visible
Jev
94.4
Forecast-only script
94.4
Fixed rules
93.1
No intervention
89.6
88909294
Final health: Jev 96.6 · Forecast-only script 96.7 · Fixed rules 95.2 · No intervention 89.8
How it works. Every 3 seconds Jev gets the city’s state and a 15-second forecast for each option, then chooses. A script that always takes the best forecast scores the same, so the forecasts do most of the work; Jev’s part is reading them, and the situation, in plain language instead of hand-coded rules.
Precisely: Jev was consistently better than fixed rules: +1.26 points on average (95% interval +1.12 to +1.40). Against fixed rules: 48 wins · 0 ties · 0 losses across 48 matched conditions. Jev was consistently better than no intervention: +4.80 points on average (95% interval +4.46 to +5.14). No reliable difference between Jev and the forecast-only script (always takes the best 15-second forecast): +0.00 points on average (95% interval −0.02 to +0.02). Jev’s briefing includes the same projections; the baseline simply picks the best one.
The score. Average health is the network health index averaged over the whole 90-second run, the complement of the game’s damage score. It rewards a fast recovery, not just a good final second. Higher is better for the operator. It is a designed congestion index, not an official transit metric.
The ceiling. In 3 of 8 scenarios, Jev and the rulebook both score between 95 and 100: the city has little room to get healthier, so differences there are fractions of a point. Operators separate in the hard scenarios (Bronx crowd, Station closure, Blackout, Compound, Late pressure), which are highlighted below.
The interval. Differences are paired: each operator faces the same seed, arrivals and disruptions. The 95% interval treats each of the 6 seeds as one observation. An interval that crosses zero means no reliable difference.
Timing. Every operator gets the same decision opportunities; the earlier protocol’s extra final-second action for fixed rules has been removed.
Prompt. metro-choice-4 is live. On these same conditions it changed average health by +1.42 points versus metro-choice-3 (nyc-jev-5 options, no projections) (95% interval +1.38 to +1.47; 47 wins · 0 ties · 1 loss).
Scope. One model run per condition (48 conditions: 6 seeds × 8 scenarios). Trains, riders and disruptions are simulated on real MTA geography; this is not real-world transit performance. “nyc-jev-6” is the version name of the current simulator (simulator v6), used for both cities.
In 3 of 8 scenarios, Jev and the rulebook both score between 95 and 100: the city has little room to get healthier, so differences there are fractions of a point. Operators separate in the hard scenarios (Bronx crowd, Station closure, Blackout, Compound, Late pressure), which are highlighted below.
Jev minus each alternative, by scenario Average network health over 90 seconds · 6 seeds per scenario · dot = mean · bar = 95% interval
972 measured Jev responses from this evaluation. Typical and slow requests both count.
Response time distribution Server round trip, including retries
0
0–100ms
0
100–200ms
387
200–300ms
294
300–500ms
206
500–1kms
85
1k+ms
327 msMedian / p50
940 ms90th percentile
1,197 ms95th percentile
1,997 ms99th percentile
$0.11Per 1,000 decisions, estimated
972 model requests · $0.10 estimated inference cost · 2,470,472 input tokens · jev-1.13.0. Inference only: hosting and failed-request usage are excluded. A fixed test sample, not a live service SLA.
Jev: 0 collapses · Fixed rules: 0 collapses · Forecast-only script: 0 collapses · No intervention: 0 collapses in 48 conditions. 0 model request errors and 0 incomplete conditions recorded. A collapse means health stayed critically low; an API error is a separate failure.
No simulated collapses in this sample. That does not establish real-world reliability.
Actual arrival predictions, service alerts and historical ridership from New York. These are separate from the simulated trains you see on Watch.
Next trains
Loading MTA arrival predictions…
Auto-refreshes every 30 seconds while this page is visible.
Service alerts
Connecting to the MTA service-alert feed…
New York, by the numbers.
427,004
subway entries in one hourSeptember 9, 2026 · 08:00–09:00 EDT 424 station complexes in the source snapshot
Grand Central-42 St (4,5,6,7,S)14,033
Times Sq-42 St/Port Authority Bus Terminal (1,2,3,7,A,C,E,N,Q,R,W,S)13,987
34 St-Penn Station (1,2,3)8,287
34 St-Penn Station (A,C,E)7,218
Jackson Hts-Roosevelt Av/74 St-Broadway (7,E,F,M,R)5,428
Flushing-Main St (7)5,357
Official MTA hourly estimates, summed across fare classes and payment methods. Historical entries, not live demand, train occupancy or exits; the simulator does not use them to calibrate demand. View the source ↗
These studies were run before the current evaluation, mostly on the previous simulator, and informed today’s prompt. Their numbers are not the current result above and are not directly comparable with it.
PROMPT STUDY / metro-choice-3 VS metro-choice-2 / NYC-JEV-5 / 2026-09-21
Did the new prompt help?
Mean final health (0–100) · 48 matched conditions · previous simulator
Jev · metro-choice-3
Jev · metro-choice-2
Fixed rules
No intervention
84.0
83.4
83.9
77.7
The newer prompt gives Jev the immediate crowd effect of each transfer. Against the previous prompt: +0.56 points, 20 better · 19 tied · 9 worse. Versus fixed rules, the seed-cluster 95% interval was −0.64 to +0.89 points, so no reliable difference. Small synthetic sample; the uncertainty interval includes no advantage. Fixed rules receives an extra terminal action under the unchanged benchmark. No claim of reliable superiority.
Original expanded evaluation · nyc-jev-5First prompt (metro-choice-1), every win, tie and failure.
ORIGINAL PROMPT / metro-choice-1 / nyc-jev-5 / jev-1.13.0 / 2026-09-21
192 trials. All outcomes.
12 seeds × 8 scenarios × 2 Jev repeats. Every trial runs for up to 90 simulated seconds with matched rules and no-intervention comparisons. All 192 trials completed; 0 trial failures. Recorded choices reproduced each Jev run’s final health and rider count exactly.
82.2/100Jev · mean final health
83.3Rules · same conditions
77.6No intervention
3,633Actual model requests
Jev improved health over no intervention by 4.5 points. The rules baseline led Jev by 1.2 points. Survival was identical on average (81.6s). This suite does not establish a Jev advantage over rules.
Mean final network health by scenario (0–100)
Scenario
Jev
Rules
No action
Quiet
98.0
98.7
96.0
Signal
97.1
98.2
89.4
Crowd at Times Square
97.9
98.7
96.0
Bronx crowd
97.3
99.8
84.3
Station closure
95.8
97.7
93.9
Blackout
95.3
95.7
93.1
Compound
16.6
16.8
13.9
Late pressure
59.2
61.3
54.5
Against rules: 20 health wins, 63 ties, 109 losses. Request latency: median 258ms, p95 391ms. Estimated model input cost: $0.242 for the suite; infrastructure excluded.
A designed simulator, not a validated transit benchmark. Repeats share conditions. Outcomes depend on the candidate actions, prompt, parameters, and 90-second horizon. Cost uses recorded input tokens and the documented rate at test time. No claim of statistical significance or real-world transit performance.
Latency includes the server’s TypeSafe round trip and any retries; it is not pure model compute time. Percentiles describe up to 1,200 recent valid responses. Cost uses documented pricing of $0.042 per million input tokens, verified September 21, 2026. Failed-request usage is unavailable; this estimate is not an invoice.
SAME CITY / THREE OPERATORS
Does Jev make the difference?
Six matched scenarios: two seeds, three disruption schedules, 90 simulated seconds each. Jev, rules, and no intervention receive the same arrivals and events. Jev and rules choose from the same valid actions every three ticks.
03
Jev. Rules. No intervention.
Run real model evaluations and inspect survival, health, latency, and token usage. No invented results.
A small simulator diagnostic, not real-world transit validation. A 90-second survival result is capped by the test window. Jev may tie or lose; all completed scenarios are shown. No API calls are made just to display previously recorded results.
DECISION LOG / CURRENT RUN
See the choice. Then the consequence.
0 of 0 matured decisions were followed by higher network health 12 ticks later. This is observational: later decisions and incident recovery also influence the result. It is not causal attribution or confidence calibration.
No model choices in this run yet.
Choose Jev as operator, disrupt a station, then return to inspect its response.
The latest eight records are shown; downloads include up to 500 recent detailed records. Replays store all applied choices. Device-local records and share links are not independently attested benchmark submissions.
Run your own baseline comparisonA local rules-only diagnostic; no Jev requests.
MEASURED HERE / RULE-BASED SIMULATION
A fair starting line.
Twelve paired scenarios. Identical seeds and disruption schedules. Compare the current rule-based operator with no intervention. This measures the simulator’s baselines—not Jev.
12
Same city. Same pressure. Two operators.
Run the comparison locally to see measured survival and final health. No sample results are prefilled.
90 simulated seconds per scenario · 3 seeds × 4 event configurations · generated in your browser · no API latency or model cost implied.
Your loaded session snapshotSeparate from the live city. Inspect or download this session.
YOUR SESSION / SPRINT / RULES
Your loaded snapshot.
100%Network health index
00:00Simulated time survived
0Operator choices
0Simulated completed trips
START OF AVAILABLE METRICS00:00 / HEALTH 0–100
This snapshot updates when you return to Watch. Personal challenges pause while you explore. Network health is a designed congestion index, not an MTA service metric.