Skip to main content
Hulash Chand
homeprojectslocus
Hulash Chand

London Transport Delay Predictor

London Transport Delay Predictor

London Transport Delay Predictor

Forecast for 23:00, Friday 2 Oct

4 lines look likely to be disrupted in an hour.

CentralMetropolitanPiccadillyVictoria

Predicted from live TfL status, recent disruption history and London weather. Updated at 22:17 using model version 2.

Underground

  • BakerlooNow: Good serviceLooks clear0%
  • CentralNow: Good serviceLikely disrupted100%
  • CircleNow: Good serviceLooks clear0%
  • DistrictNow: Good serviceLooks clear0%
  • Hammersmith & CityNow: Good serviceLooks clear0%
  • JubileeNow: Good serviceLooks clear0%
  • MetropolitanNow: Good serviceLikely disrupted100%
  • NorthernNow: Good serviceLooks clear0%
  • PiccadillyNow: Good serviceLikely disrupted100%
  • VictoriaNow: Severe delaysLikely disrupted100%
  • Waterloo & CityNow: Good serviceLooks clear0%

Elizabeth line

  • Elizabeth lineNow: Good serviceLooks clear0%

DLR

  • DLRNow: Good serviceLooks clear0%

Overground

  • LibertyNow: Good serviceLooks clear0%
  • LionessNow: Good serviceLooks clear0%
  • MildmayNow: Minor disruptionLooks clear0%
  • SuffragetteNow: Good serviceLooks clear0%
  • WeaverNow: Good serviceLooks clear0%
  • WindrushNow: Good serviceLooks clear0%

In the last 24 hours it caught 106 of 172 disruptions, with 144 false alarms across 1430 checked forecasts.

Part 1 of 3 end-to-end Databricks projects: live TfL line status and London weather are ingested every 15 minutes, refined through a Bronze to Silver to Gold Lakeflow pipeline, and turned into a model that predicts whether a line will be disrupted one hour from now. Built entirely on Databricks Free Edition (serverless only).

15 minIngestion Interval
3Lakeflow Layers
1 hourForecast Horizon
Free EditionServerless Only

Live Forecast

Reads the same public snapshot the Databricks job pushes to GitHub every hour. Cached server-side for 15 minutes, so this loads instantly rather than waiting on a client-side fetch.

Architecture

A scheduled Lakeflow pipeline, not a continuous stream: Free Edition has daily compute quotas and a cap on concurrent serverless resources, so ingestion polls every 15 minutes and Auto Loader processes only new files.

01

Ingest

Every 15 min

A scheduled job polls the TfL Unified API for live line status and Open-Meteo for London weather, landing the raw JSON in a Unity Catalog volume.

TfL Unified APIOpen-Meteo APIUC Volume
02

Bronze

Auto Loader

Auto Loader picks up new files only and lands them untouched, full API payload plus file lineage columns, so Silver can be rebuilt at any time.

line_status_rawweather_raw
03

Silver

Hourly at :07

Flattened, labelled and deduplicated. Grain is a status, not a line, since a line can carry two statuses at once, e.g. part closure plus minor delays.

line_statusweather
04

Gold

Feature table

One row per line per 15-minute slot, with the 1-hour-ahead disruption target attached. Planned engineering closures are excluded from the target and flagged separately instead.

line_features
05

Model + App

MLflow

Batch scoring writes predictions back to Gold, and a Databricks App publishes a public snapshot to GitHub every hour, the same snapshot the widget above reads.

MLflow scoringDatabricks AppPublic snapshot

Design Decisions

01Cost control

Scheduled micro-batches, not a continuous stream

Free Edition has daily compute quotas and a cap on concurrent serverless resources. Ingestion polls every 15 minutes and Auto Loader processes only new files: the same incremental semantics as streaming, at a fraction of the cost.

02Scheduling

Ingestion and pipeline run on separate schedules

Ingest every 15 minutes (cheap), pipeline hourly at :07, so the two never compete for compute.

03Data modelling

Bronze stays raw, Silver grain is a status

The full API payload is kept untouched in Bronze with file lineage columns. Silver is graded per status rather than per line, since a line can legitimately carry two statuses at once.

04Label quality

Planned works are excluded from the target

Scheduled engineering closures aren't something you predict from live conditions. They're flagged with is_planned instead of polluting the disruption target.

05Correctness

Time features use London local time

So the BST/GMT switch doesn't silently shift the hour of day a row is attributed to.

Gotchas I Hit

Windows CLI keyring error ("OS keyring unreachable"): fixed with DATABRICKS_AUTH_STORAGE=plaintext before logging in, then passing a profile flag to every command.

TfL returned 429 "Invalid app_key": the stored key was 33 characters because a pasted null character had snuck in. Checking the key's length found it in seconds.

RESOURCE_EXHAUSTED when starting the SQL warehouse: that's the concurrent serverless limit on Free Edition. Ad-hoc queries now run from a notebook already on serverless instead.

DatabricksLakeflowDelta LakeAuto LoaderMLflowUnity CatalogPySparkDatabricks Asset BundlesTfL Unified APIOpen-Meteo APIDatabricks Apps

Explore the pipeline, the notebooks, and the setup instructions:

View on GitHub
Back to all projects

Let's build something.

Now Playing

Utopia

Horacio Pagani