Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Uber’s 2017 approach to unusual ride-demand spikes was not simply “use an LSTM.” It trained one recurrent neural network across thousands of time series from multiple cities, supplied it with historical demand and external signals, then added an automatic feature-extraction module after a vanilla shared LSTM failed to beat the baseline. Uber reported improvements against three different benchmarks—but those figures describe a historical case study, not a current specification of Uber’s forecasting system.
The forecasting problem: where, when and how many
For operations, a demand forecast needs to estimate where ride requests will occur, when they will arrive and how many to expect. That informs resource allocation, planning, anomaly detection and budgeting. Forecast errors become especially consequential around periods when demand changes abruptly: New Year’s Eve and New Year’s Day, Christmas, concerts, sporting events, inclement weather and other local events.
In Uber’s June 9, 2017 engineering account, “extreme event” means an unusual operational demand period; it does not refer to a formally defined extreme-value-theory model. The difficulty is that these periods matter precisely because they can differ from an ordinary day, while the historical record may contain only a few comparable examples. An annual holiday such as New Year’s Eve recurs, but it still supplies very few independent observations for any one city or area.
Nor is a holiday a fixed input with a fixed effect. Population growth, marketing changes, weather, driver incentives and local events can all shift demand. Cities also differ in scale, seasonality and response to the same event. Underpredicting a major peak can leave operations short of supply; overpredicting can waste resources. Forecast quality therefore depends on more than recognizing a date on the calendar.
#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
Why one model for many time series?
Uber described a forecasting environment with many metrics and city-level series. Fitting an independent model for each series can be difficult when individual histories are sparse and behavior changes. A model trained across many cities and thousands of series offers a way to share statistical strength: patterns learned from one series may help with another whose own history is limited.
That does not mean treating every city as identical. A shared model must still recognize which kind of series it is forecasting and how its context differs. This became the central engineering issue in Uber’s account: the first, vanilla LSTM did not outperform the existing baseline. Uber said the model did not adapt well to time-series domains absent from training and did not distinguish heterogeneous series sufficiently. Manually providing identifying features for millions of metrics was not a practical solution.
What an LSTM offered—and what it needed
An LSTM is a recurrent neural network built to carry and update information across a sequence. Uber’s rationale included end-to-end modeling, automatic feature extraction, and the ability to incorporate external variables and learn nonlinear relationships among many dimensions. Those properties made an LSTM an attractive candidate for a large forecasting system, but they did not make a generic shared LSTM automatically suitable for every city and series.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
The reported response was a custom, ensemble-based feature-extraction module. At a high level, the system generated feature vectors, averaged the extracted vectors using an ensemble technique, concatenated that representation with the model input, and used the combined representation to produce a forecast. Conceptually:
Rank #2
Historical demand + available external signals
↓
Automatic feature extraction
↓
Average extracted feature vectors
↓
Concatenate features with model input
↓
LSTM forecast
This is a conceptual summary, not a reproducible model diagram. Uber’s public account does not disclose enough architectural detail to recreate the feature extractor or full network exactly. The core lesson is narrower and more useful than “LSTMs work”: a global model needs a way to represent differences among the series it shares.
Inputs, preparation and sliding windows
The model drew on scaled trip counts over time and external information. Uber named precipitation, wind speed and temperature forecasts; trips in progress within a geographic area; registered users; local holidays and events; and other city-level information. These inputs are intended to help the model distinguish, for example, a demand change associated with weather from one associated with a recurring calendar pattern.
The article names log transformation, scaling and detrending as preprocessing steps. It does not specify the exact formulas, scaling method, missing-data treatment, feature frequency or leakage controls, so those details should not be inferred.
Free tools Windows power users keep installed
One-click scans. No signup required.
Training used the standard sliding-window idea for supervised sequence forecasting. An input window X contains a fixed run of historical time steps and features; the target window Y contains the future values to predict. The windows shift forward through the history, producing many training examples. The source describes tensor dimensions conceptually as batch, time and features for X, with forecasted features for Y, and gives mean squared error as an example loss. It does not publish every window dimension, lag length, optimizer or hyperparameter.
One important practical condition follows from using external signals: each feature must be available at the time a forecast is issued. A valid historical weather input is the forecast then available, not the weather that was later observed. The article does not document how Uber handled this issue, so it remains a methodological detail that cannot be assumed from the public description.
What the holiday example tested
Uber illustrated the work with five years of daily completed-trip history across U.S. cities. The example examined a seven-day interval before, during and after major holidays, including Christmas Day and New Year’s Day. In that experiment, Christmas Day was among the most difficult to predict, with the greatest error and uncertainty in rider demand described by the authors. That is a result for the reported experiment, not a universal ranking of holidays across markets or years.
Daily totals also set a limit on what the example demonstrates. A forecast that captures daily volume need not capture the timing of an hourly or sub-hourly surge. The published holiday example should not be treated as evidence of performance at every operational time scale.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Three reported comparisons, three different baselines
| Comparison in Uber’s report | Reported result |
|---|---|
| Custom architecture versus base LSTM | 14.09% SMAPE improvement |
| Custom architecture versus the classical time-series model used in Argos | More than 25% improvement |
| New model versus Uber’s prior proprietary model in the described testing | 2–18% accuracy increase |
SMAPE is a percentage-based error metric. The 14.09% figure is specifically reported against the base LSTM; it is not the overall production gain. The more-than-25% figure uses the classical model in Argos, Uber’s real-time monitoring and root-cause-exploration tool, as its comparator. The 2–18% figure is described as an accuracy increase against a proprietary model. It should not be recast as a 2–18% reduction in error: the article does not establish that equivalence.
Rank #4
These values cannot be combined into a single headline improvement. A percentage comparison is difficult to interpret fully without the underlying baseline error, the evaluated series and period, aggregation method, out-of-sample design and precise definition of “accuracy.” Uber’s public account does not provide confidence intervals, statistical significance, per-city results or a complete error distribution. The numbers are Uber-reported results, not independently reproduced benchmarks.
From offline training to inference
Uber said it trained the network offline with TensorFlow and Keras, exported the learned weights, and implemented inference in native Go. Separating training from serving can keep a production inference path independent of the heavier training stack. It also makes parity important: teams adopting this pattern need to check that the exported model produces acceptably equivalent outputs in the serving implementation, preserve reproducible model artifacts, and test updates before rollout. These are general engineering considerations, not failure reports about Uber’s system.
The 2017 article says the model was used in production, but it does not disclose the deployment topology, number of markets, inference latency, forecast refresh cadence, monitoring thresholds, rollback process, human overrides, serving cost or current system status. It is accurate to describe what Uber reported then; it is not evidence that Uber uses this exact architecture today.
Recommended Free Tools
When a similar global neural model makes sense
Uber’s own decision criteria are a useful starting point: neural networks are more likely to help when there are many time series, the series are long, and meaningful correlations exist among them. A pooled LSTM-style approach is more plausible when individual histories are sparse but related series can contribute useful information, external signals are available at forecast time, and a team can maintain the shared data and model pipeline.
Best Value
It may be a poor fit when there are only a few short series, series are weakly related or generated by fundamentally different processes, external inputs are unreliable, or rare events cannot be evaluated robustly. A neural model also brings operational complexity that may not pay off against strong statistical baselines. Classical methods can remain effective for short, stable, well-understood series; gradient-boosted or probabilistic models may be better choices under other data and operational constraints. The 2017 account does not compare Uber’s model with current transformer, gradient-boosting, probabilistic or foundation-model forecasters.
- Global versus local: A global model shares information and can simplify management across many series. Local models preserve independence and may fit specialized behavior better. Pooling risks underfitting unusual series or transferring patterns poorly between cities.
- Point accuracy versus operational cost: SMAPE may not reflect the cost of a missed peak, and percentage metrics can behave awkwardly near very small actual values. Consider absolute or weighted error, peak underprediction, service-level impact and other decision-relevant measures.
- Forecast versus uncertainty: The article discusses error and uncertainty around holidays but does not describe a probabilistic output head or calibrated prediction intervals. Do not assume a point-forecast LSTM provides reliable uncertainty estimates.
Validation and failure modes to plan for
Rare events make evaluation especially easy to get wrong. Randomly splitting individual observations can put near-identical periods into training and test sets, overstating performance. For a modern implementation, use rolling-origin backtests and, where data permits, hold out entire events or years. Check performance by city and event type, not just an aggregate score. These are recommended evaluation practices, not methods Uber confirms in its article.
Also test for event drift. Population, service area, prices, incentives, marketing and customer behavior can change the response to the same holiday. Weather forecasts and event schedules must be represented as they were known at issuance time. Pooling cities can cause negative transfer when calendars, demand scales, data quality or seasonal patterns differ. Finally, daily data may conceal intraday peaks, and a model’s success on one horizon does not establish success at another.
What the public account leaves unknown
Uber’s article is an engineering account, not a complete reproducible research paper. It does not publish the exact network layers and sizes, optimizer, regularization, training schedule, complete feature list, backtesting protocol, confidence intervals or full benchmark tables. It also does not specify model export tooling, serving latency, retraining cadence or present-day production status. Those are real limits on reproduction and comparison—not details to fill in by guesswork.
The enduring contribution is the system-design lesson: for sparse, heterogeneous demand series, sharing a model can help, but only if the model receives enough information to distinguish the series and their contexts. Uber’s reported results support the promise of that tailored 2017 approach; they do not prove that LSTMs outperform classical forecasting everywhere or that the same architecture remains the right choice today. The primary account is Uber Engineering’s 2017 article.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

