London Transport Delay Predictor
Forecast for 23:00, Friday 2 Oct
4 lines look likely to be disrupted in an hour.
Underground
- BakerlooNow: Good serviceLooks clear0%
- CentralNow: Good serviceLikely disrupted100%
- CircleNow: Good serviceLooks clear0%
- DistrictNow: Good serviceLooks clear0%
- Hammersmith & CityNow: Good serviceLooks clear0%
- JubileeNow: Good serviceLooks clear0%
- MetropolitanNow: Good serviceLikely disrupted100%
- NorthernNow: Good serviceLooks clear0%
- PiccadillyNow: Good serviceLikely disrupted100%
- VictoriaNow: Severe delaysLikely disrupted100%
- Waterloo & CityNow: Good serviceLooks clear0%
Elizabeth line
- Elizabeth lineNow: Good serviceLooks clear0%
DLR
- DLRNow: Good serviceLooks clear0%
Overground
- LibertyNow: Good serviceLooks clear0%
- LionessNow: Good serviceLooks clear0%
- MildmayNow: Minor disruptionLooks clear0%
- SuffragetteNow: Good serviceLooks clear0%
- WeaverNow: Good serviceLooks clear0%
- WindrushNow: Good serviceLooks clear0%
In the last 24 hours it caught 106 of 172 disruptions, with 144 false alarms across 1430 checked forecasts.
Part 1 of 3 end-to-end Databricks projects: live TfL line status and London weather are ingested every 15 minutes, refined through a Bronze to Silver to Gold Lakeflow pipeline, and turned into a model that predicts whether a line will be disrupted one hour from now. Built entirely on Databricks Free Edition (serverless only).
Live Forecast
Reads the same public snapshot the Databricks job pushes to GitHub every hour. Cached server-side for 15 minutes, so this loads instantly rather than waiting on a client-side fetch.
Architecture
A scheduled Lakeflow pipeline, not a continuous stream: Free Edition has daily compute quotas and a cap on concurrent serverless resources, so ingestion polls every 15 minutes and Auto Loader processes only new files.
Ingest
Every 15 minA scheduled job polls the TfL Unified API for live line status and Open-Meteo for London weather, landing the raw JSON in a Unity Catalog volume.
Bronze
Auto LoaderAuto Loader picks up new files only and lands them untouched, full API payload plus file lineage columns, so Silver can be rebuilt at any time.
Silver
Hourly at :07Flattened, labelled and deduplicated. Grain is a status, not a line, since a line can carry two statuses at once, e.g. part closure plus minor delays.
Gold
Feature tableOne row per line per 15-minute slot, with the 1-hour-ahead disruption target attached. Planned engineering closures are excluded from the target and flagged separately instead.
Model + App
MLflowBatch scoring writes predictions back to Gold, and a Databricks App publishes a public snapshot to GitHub every hour, the same snapshot the widget above reads.
Design Decisions
Scheduled micro-batches, not a continuous stream
Free Edition has daily compute quotas and a cap on concurrent serverless resources. Ingestion polls every 15 minutes and Auto Loader processes only new files: the same incremental semantics as streaming, at a fraction of the cost.
Ingestion and pipeline run on separate schedules
Ingest every 15 minutes (cheap), pipeline hourly at :07, so the two never compete for compute.
Bronze stays raw, Silver grain is a status
The full API payload is kept untouched in Bronze with file lineage columns. Silver is graded per status rather than per line, since a line can legitimately carry two statuses at once.
Planned works are excluded from the target
Scheduled engineering closures aren't something you predict from live conditions. They're flagged with is_planned instead of polluting the disruption target.
Time features use London local time
So the BST/GMT switch doesn't silently shift the hour of day a row is attributed to.
Gotchas I Hit
Windows CLI keyring error ("OS keyring unreachable"): fixed with DATABRICKS_AUTH_STORAGE=plaintext before logging in, then passing a profile flag to every command.
TfL returned 429 "Invalid app_key": the stored key was 33 characters because a pasted null character had snuck in. Checking the key's length found it in seconds.
RESOURCE_EXHAUSTED when starting the SQL warehouse: that's the concurrent serverless limit on Free Edition. Ad-hoc queries now run from a notebook already on serverless instead.
Explore the pipeline, the notebooks, and the setup instructions:
View on GitHub