RUN BY JEVPlay the daily ↗

METHODOLOGY / OPEN ASSUMPTIONS

Trust the record.
Question the model.

A credible experiment separates observed infrastructure, simulated behavior, and measured operator performance.

What is real?

Live arrival predictions & service alertsMTA GTFS-Realtime and TfL arrival predictions and line status, with freshness and missing-feed states. TfL line status is labeled with our check time. Displayed as context; they do not drive simulated operations.OBSERVED / LIVE
Station locations & service connectionsOfficial MTA regular GTFS. Weekday northbound patterns; some short turns omitted. Platforms merged using transfer records. London uses TfL ordered Tube route sequences, including branches, for 272 stations and 11 lines.OBSERVED DATA
Map geometryActual station coordinates in both cities, with NYC boundary polygons and London borough/Thames polygons. Edges join scheduled stops; they are not surveyed track paths.SIMPLIFIED DISPLAY
Historical subway entriesOfficial MTA hourly entry estimates, September 9, 2026, 08:00–09:00 EDT. London uses TfL NUMBAT 2025 typical weekday boardings, including interchanges—not unique passengers. Both are historical references only, not demand calibration.OBSERVED DATA
Demand, crowding & train operationsSimulated passenger arrivals, train capacity, travel times, and incidents. No origin–destination assignment or timetable validation yet.SIMULATED
OperatorJev selects from code-validated actions every three ticks; a rule-based comparator uses the same options. Nine initial reserves. API outputs, confidence, latency, and token usage are recorded.MODEL + CODE

What the numbers mean.

NETWORK HEALTH

100 minus a weighted penalty from the 14 most crowded complexes and the fraction of held trains, clamped to 0–100. This is a designed game index.

RECORDED TIME

Elapsed simulated seconds until health stays below 22 for seven ticks. Ongoing mode has no timer cutoff. Sprints end at 90 seconds.

RIDERS MOVED

Simulated alightings, counted as completed trips. They are not unique people or actual transit ridership.

How comparisons stay fair.

THE HEADLINE RESULT / SIMULATOR V6

New York

Yes. Jev ran New York’s subway better than a purpose-built rulebook. +1.3 health points on average, and ahead in 48 of 48 matched tests. Each decision takes about 0.3 s and costs $0.0001.

Average health: Jev 94.4 · rules 93.1 · none 89.6. Jev − rules +1.26 (95% interval +1.12 to +1.40; 48 wins · 0 ties · 0 losses). Jev − none +4.80. 48 conditions · metro-choice-4 · 2026-09-22 · full evidence ↗

London

Yes, by a small margin: Jev ran London’s Underground a little better than a purpose-built rulebook. +0.5 health points on average, and ahead in 48 of 48 matched tests. Each decision takes about 0.3 s and costs $0.0001.

Average health: Jev 88.2 · rules 87.7 · none 86.0. Jev − rules +0.51 (95% interval +0.35 to +0.66; 48 wins · 0 ties · 0 losses). Jev − none +2.23. 48 conditions · metro-choice-4 · 2026-09-22 · full evidence ↗

Paired baselines receive the same network, seed, arrival randomness and event schedule; only the intervention policy changes. The current evaluation (simulator v6, version name “nyc-jev-6”, used for both cities) runs 48 matched conditions per city: 6 held-out seeds × 8 scenarios, one Jev run per condition, compared with fixed rules, a best-15-second-projection baseline and no intervention. The primary metric is average network health across the 90-second run. Intervals treat each seed as one observation. Every operator gets the same decision opportunities; the earlier protocol’s extra final-second action for fixed rules was removed.

Research history. Earlier studies used the previous simulator, nyc-jev-5, before player disruptions were rebalanced and Jev’s options were widened; demand and the health index are unchanged, so those results describe an earlier version of the game. The metro-choice-3 prompt study compared 48 matched conditions; mean final health was 83.98 for metro-choice-3 and 83.85 for rules, with a seed-cluster interval crossing zero. The prompt was kept despite not meeting the predefined reliability threshold. The original expanded evaluation used 12 seeds, 8 scenarios and 2 model repeats (192 Jev trials); repeats share conditions and are not independent. London is evaluated separately with 48 matched conditions. The interactive six-scenario check and the rules-only diagnostic are small diagnostics, not evidence of broad real-world generalization.

NETWORK nyc-gtfs-20260826-v1FEED 20260826-X-long-term-supplement-trip-idsSIMULATION nyc-jev-6POLICY jev / rules / noneJEV VERSION jev-1.13.0CURRENT EVALUATION nyc-jev-6

Replay, performance, and limits.

Short links reference replay inputs saved on the server: version, seed, mode, disruption schedule, and applied model choices. Earlier self-contained links remain supported. Anyone with the link can open the run; no account is needed. Replay applies recorded choices without new API calls. Model output itself is not assumed deterministic; a rematch calls Jev again. Results are reconstructed from the saved actions. Unranked shared runs reconstruct the outcome but do not authenticate submitted model actions. Ranked daily and friend challenges run on the server, which calls Jev, validates moves and derives the score. Standings represent anonymous browser guests, not verified people. A worker handles replay and baseline computation. The display retains at most 600 detailed frames, metrics are downsampled, and the last 30 run records stay on this device. Replay links are capped at 24 simulated hours and 5,000 events to bound computation.

Watch uses a shared server-operated world saved between visits. One model decision serves all viewers. The deployed world advances three simulated seconds per scheduled update, while viewed, pausing after two minutes without viewers to limit costs. Updates are scheduled after the previous decision completes, so simulation time is not wall-clock time. Saved state survives restarts. A new round begins after collapse or 15 simulated minutes. Challenge Jev copies the exact current state into a personal 90-second run. Personal challenges pause when hidden or when browsing other pages. Live MTA and TfL feeds are displayed separately as context; they do not change simulated demand or train operations. Service failures pause Jev runs with a retry option; they do not silently switch to rules.

Guest challenges and standings

Daily challenges reset at midnight UTC, separately for London and New York. Friend standings use the same starting conditions as the shared challenge. Jev choices and scores are recorded by the server; its responses may differ between attempts. Earlier collapse wins, otherwise the most damage wins: the average health lost across the 90 seconds. Each move can be used once, and each daily challenge rules out one move for the day. Ties share a rank. Each guest browser keeps one personal best per board. Clearing storage creates a new guest, so entry counts are not unique people.

A random guest identifier and optional nickname are stored in your browser. The server stores a hash of the guest identifier to associate your scores. Saving or sharing publishes your nickname, result and replay. Tab-scoped browser storage keeps a run credential and pending moves so a refresh can recover your challenge from the server. Runs expire for play after two hours; published results remain available. Share-card ranks are snapshots at submission. Guest names are unverified. Unranked play remains available.

Simulation research records

Completed shared rounds and server-run ranked challenges can be retained privately for up to 90 days to improve the simulation controller. Records contain city and software versions, random seeds, disruptions, AI choices, bounded decision telemetry and outcomes. They exclude guest identifiers, nicknames, IP addresses and credentials. Collection is capped at 500 records per day and 64 KiB per record; capped or unavailable archives can leave gaps. These records do not automatically train or change Jev. Candidate changes are evaluated separately before release.

Anonymous product measurement

We also use Usermaven in cookie-less mode to measure page views and game actions, such as challenge completion and sharing. Analytics requests go directly to Usermaven. We do not send nicknames, guest identifiers, run credentials or replay contents as custom event properties. Automatic click and form tracking are enabled.

We count challenge starts, completions, replay opens and shares. A random identifier lasts for your browser tab session and connects a shared link to subsequent play and sharing. These counts include testing and do not identify unique people. The activity report shows observed sharing and play-through ratios with their sample sizes; a shared link is attributed for up to seven days, so a journey that crosses midnight UTC appears in both days; known creator tab sessions are excluded from shared-link cohorts. Other self-opens, blocked requests and small samples can still distort these observations. We also collect loading, responsiveness and layout stability metrics for a small set of public pages, without query strings or personal details. Reports are bounded daily samples; small samples do not establish retention, virality or search performance.