Revisiting My 2020 LSTM Stock-Prediction Project: From Pretty Charts to Honest Forecasting

Revisiting My 2020 LSTM Stock-Prediction Project: From Pretty Charts to Honest Forecasting

Six years ago I wrote an article called “Predicting stock prices using a TensorFlow LSTM neural network for time-series forecasting.”

At the time, the project was exactly what I wanted it to be: a practical experiment to understand how Long Short-Term Memory networks could be applied to financial time series.

I downloaded market data from Yahoo Finance, normalized closing prices with MinMaxScaler, converted them into rolling sequences, trained a stacked LSTM network and compared the predicted price series against the real one.

The resulting charts looked surprisingly convincing.


Since then, the repository has continued to attract users, forks and stars, and I recently decided to go back to it properly.

Not simply to update TensorFlow dependencies or make the code prettier, but to answer a much harder question:

Was the model actually learning useful forecasting information, or was it simply producing a visually convincing approximation of an extremely persistent price series?

That question changed the direction of the project.

The repository has now evolved from a single LSTM stock-prediction experiment into what I would describe as a small quantitative forecasting laboratory.

The problem with a good-looking stock chart

Consider the following sequence:

Actual price:

650 → 653 → 651 → 654 → 655

A neural network might predict:

650 → 652 → 650 → 653 → 654

Plot those two lines and the result looks excellent.

But there is a very important baseline:

Tomorrow's price = today's price

That gives:

650 → 650 → 653 → 651 → 654

Financial price levels are highly persistent.

This means a model can produce an impressive chart, achieve a low RMSE and still be worse than simply predicting that tomorrow's closing price will equal today's.

That is the first major lesson from revisiting the project.

Absolute model error is not enough.

The current project therefore calculates skill relative to the naïve forecast:

RMSE skill = 1 - RMSE(model) / RMSE(naive)

If the value is positive, the model beats the naïve previous-price forecast.

If it is negative, the sophisticated neural network lost to:

tomorrow = today

That turned out to be a very useful reality check.

Rebuilding the experiment around validation

The second major change was methodological.

My original experiment was designed primarily as an implementation exercise. Like many early machine-learning examples, the distinction between validation and final test data was not strong enough.

The updated training pipeline now gives the three datasets very specific responsibilities:

TRAIN
  │
  │ optimise model weights
  ▼
VALIDATION
  │
  │ early stopping / learning rate
  ▼
TEST
  │
  │ final evaluation only
  ▼
RESULT

The test period is no longer involved in deciding when training should stop.

For research comparisons I also added chronological walk-forward evaluation so models can be tested across multiple market regimes rather than one favourable period.

Random train/test shuffling is particularly dangerous for financial time series because the future must not leak backwards into the past.

The project now treats chronological validation as a first-class requirement.

v7 versus v8: does a shared representation help?

The existing project had evolved through several generations of LSTM models.

One of the more interesting versions was v7.

v7 predicts two things:

LSTM 1 ──> direction

LSTM 2 ──> magnitude

The direction network estimates whether the next movement is positive or negative, while the magnitude network estimates the size of that movement.

For v8 I changed this architecture.

Instead of training two independent LSTMs, both tasks share the same temporal representation:

                     ┌──> direction
                     │
input ─> LSTM ─> LSTM┤
                     │
                     └──> magnitude

The hypothesis was simple: direction and magnitude are related problems, so forcing them to learn independently may be wasting useful information.

Rather than judge this from one chart, I ran a seeded walk-forward benchmark.

The test covered:

SPY
QQQ
AAPL
FTSE 100
BTC-USD

across four different regimes:

2020
2022
2024
2026 YTD

That produced 20 asset/year test slices. Their training histories overlap, and several assets are correlated, so they should not be treated as 20 independent experiments.

The result was interesting.

Measurementv7v8
Folds beating the other architecture317
Folds beating naïve baseline03
Mean RMSE skill-19.37%-7.44%
Median RMSE skill-14.79%-5.42%
Mean directional accuracy52.62%52.47%
Mean delta information coefficient-0.034+0.017

The committed fold table lets readers recompute the win counts and RMSE-skill summaries. The directional and information-coefficient figures come from the original run report; its raw market and prediction snapshots were workflow artifacts and are not retained in this checkout.

In this benchmark, v8 had lower RMSE than v7 more often.

It beat v7 on RMSE skill in 17 of the 20 comparisons.

But there was a much more important result:

v8 still failed to beat the naïve forecast consistently.

Its only three positive-skill folds were QQQ 2020, BTC-USD 2020 and SPY 2020, and even those improvements were small.

That is exactly the kind of result that disappears if the only output we inspect is a predicted-price chart.

Changing the question: v9

At this point I decided not to spend another round simply changing the number of LSTM units.

The target itself needed reconsidering.

Instead of asking:

What will the FTSE price be tomorrow?

v9 asks:

What is the next-session return, what is the probability of an upward move, and how uncertain is that forecast?

The model now operates on log returns rather than directly optimizing the next absolute price level.

Conceptually:

                     LAST 60 RETURNS
                            │
                            ▼
                       LSTM 128
                            │
                            ▼
                        LSTM 64
                            │
              ┌─────────────┼─────────────┐
              │             │             │
              ▼             ▼             ▼
          Direction      Expected       Quantile
            head          return          head
              │             │             │
            P(up)         E[r+1]      q10/q50/q90
              │             │             │
              └─────────────┼─────────────┘
                            │
                            ▼
                    held-out evaluation

The current TensorFlow model has approximately 116,000 trainable parameters.

The uncertainty head is also constructed so that its quantiles remain ordered.

Rather than allowing three independent outputs to cross each other, the model predicts a median and positive upper/lower distances:

q10 = median - lower_distance
q50 = median
q90 = median + upper_distance

This guarantees:

q10 <= q50 <= q90

by construction.

Keeping v9 inside the original project

One requirement I had when implementing v9 was that it should not become a completely separate research notebook.

The notebook remains intentionally thin.

It still uses the original project abstractions:

from stock_prediction_class import StockPrediction
from stock_prediction_lstm import LongShortTermMemory
from stock_prediction_numpy import StockData
from stock_prediction_plotter import Plotter
from stock_prediction_deep_learning import train_LSTM_network
from stock_prediction_deep_learning_inference import InferenceRunner

Training v9 is therefore just another configuration of the same framework:

train_LSTM_network(
    stock_prediction,
    use_returns=True,
    model_version="v9",
    forecast_horizon=1,
    validation_fraction=0.15,
    seed=42,
)

StockData owns the transformation.

LongShortTermMemory owns the architecture.

train_LSTM_network owns training and evaluation.

Plotter owns the visualizations.

InferenceRunner owns future forecasting.

That was important to me because the objective was to evolve the existing project, not replace it with an unrelated notebook.

What did v9 actually achieve?

The executed example from 27 September 2026 uses the FTSE 100 and evaluates 492 held-out sessions.

The model uses a 60-session input window and was trained with a maximum of 20 epochs, fixed seed 42 and chronological validation.

Training stopped after 10 of the planned 20 epochs and restored the best validation weights.

The held-out results were:

Metricv9
Return RMSE0.008549
Return MAE0.006453
RMSE skill vs zero-return-14.30%
Return-sign accuracy44.51%
Direction-head accuracy55.49%
Always-up baseline accuracy55.49%
Return information coefficient-0.0537
q10-q90 empirical coverage93.09%
Reconstructed-price RMSE skill vs naïve-14.52%

This is not a winning forecasting model.

And that is a useful result.

The expected-return head performs worse than predicting zero return on this particular held-out period.

Its correlation with realized next-session returns is also slightly negative.

The dedicated direction head gets approximately 55.49% of directions right, but a closer audit changes the interpretation: it predicts up on every held-out window. Up moves account for 55.49% of this test period, so the head only matches an always-up baseline. It does not demonstrate a useful classification signal.

This collapse gives me a concrete question for the next experiment: can a direction model beat a constant-class baseline on untouched periods, and can the return head improve on zero return?


The lines track closely in price space, but the naïve previous-close forecast still has lower RMSE. The chart alone cannot establish predictive skill.

The uncertainty result is interesting too

The q10-q90 interval is nominally intended to represent an 80% range.

On the held-out FTSE sample its empirical coverage was approximately:

93.09%

The quantile estimates are not calibrated to their nominal levels. The empirical proportions of realized returns below q10, q50 and q90 were approximately 1.02%, 21.14% and 94.11%, rather than 10%, 50% and 90%. The excess central coverage alone does not establish that the interval is simply too wide; it may also be miscentered or asymmetric.

That empirical check is more informative than simply plotting random future trajectories and calling them confidence intervals.

The project now makes a distinction between:

scenario paths

and nominal one-step quantile estimates whose calibration is measured on held-out data.

They are not the same thing.

Recursive forecasts have another problem

There is another issue that appears once we stop predicting one day ahead and start recursively forecasting 30 days into the future.

Suppose the network predicts a small negative return.

That prediction becomes part of the input for the next prediction.

The second synthetic observation becomes input to the third.

Then the third feeds the fourth.

After enough steps, the network is no longer operating primarily on observed market history.

It is operating on its own predictions.

Small bias can compound quickly.

In one FTSE run, the raw recursive v9 return remained around -0.4% per session. Compounding that mechanically for 30 sessions creates an implausibly aggressive downward trajectory.

So I added an explicit long-horizon return anchor.

The first forecast remains untouched.

After that, the model contribution decays exponentially towards the mean return observed in the fit portion of the training period.

For the current configuration the model weight has a five-session half-life:

Forecast session     Model weight
---------------------------------
1                    100%
6                     50%
11                    25%
21                   6.25%
30                   ~1.8%

For the executed FTSE example:

Latest FTSE close:     10,695.30
First predicted move:      -0.37%

The unanchored recursive path ended near:

9,390

while the anchored path ended around:

10,389

The important point is that the anchor does not improve the held-out one-step test metrics.

It exists only to make recursive long-horizon inference less dominated by accumulated model bias.

Both paths are written to the CSV so the difference remains visible rather than being hidden.

The shaded scenario and one-step residual bands are diagnostics; the latter has no validated 30-session coverage guarantee.

Building the laboratory around the model

The model itself is now only one part of the repository.

The project has also gained a maintained research layer for:

  • naïve-baseline evaluation;
  • RMSE and MAE skill calculations;
  • expanding-window validation;
  • directional accuracy and information coefficient;
  • nominal return quantiles and one-step residual bands;
  • transaction-cost-aware diagnostics;
  • recorded random seeds, with numerical reproducibility still dependent on runtime and hardware;
  • experiment metadata and SHA-256 data hashes;
  • cross-asset benchmarking;
  • automated tests and GitHub Actions;
  • CI-executed notebooks with saved artifacts.

New training runs store their configuration, market-data snapshot, software versions and Git revision alongside the trained model. Historical benchmark input snapshots were kept as workflow artifacts rather than in this checkout, so that older run cannot be reproduced byte-for-byte from the committed fold table alone.

The objective is to leave enough provenance to audit an experiment later. Exact numerical replay also depends on retaining the input snapshot and matching the runtime.

One surprising lesson from the benchmark

I also added a very simple trading diagnostic based on the direction of each prediction.

This exposed another potential trap.

A model can produce a respectable-looking Sharpe ratio simply by remaining almost permanently long during a rising market.

That does not mean the forecasting model discovered a useful trading signal.

It may simply have recreated long-only exposure.

This is why the current project separates:

forecast accuracy

from:

strategy diagnostics

and neither one is considered sufficient on its own.

A backtest is only as credible as the assumptions behind it: transaction costs, turnover, financing, borrow, slippage, capacity and market impact all matter.

The small diagnostic in this project deliberately does not pretend to be an execution simulator.

What I learned by revisiting my own work

The most useful part of revisiting a project after six years is seeing how much the definition of “working” changes.

In 2020 I was excited that the LSTM could reproduce the shape of a stock-price series.

I still think that was a useful experiment.

It taught me a lot about TensorFlow, LSTMs, time-series preparation and model training.

But today I would ask harder questions:

Did it beat the simplest sensible baseline?

Was the test data genuinely untouched?

Did the result repeat across different regimes?

Are preprocessing statistics fit using future information?

Does the model predict direction or merely price persistence?

Are uncertainty intervals calibrated?

Does an apparent trading result survive transaction costs?

Can somebody reproduce the experiment exactly?

Those questions are now more important to me than whether two lines overlap nicely on a chart.

Running v9

The original notebook workflow is still available, and the new v9 notebook follows the same structure.

After installing the dependencies:

python -m pip install -r requirements.txt

open:

examples/notebooks/stock_prediction_lstm_v9.ipynb

The notebook will:

download the market data
        ↓
prepare the return sequences
        ↓
train the v9 multi-task LSTM
        ↓
evaluate the held-out period
        ↓
generate the project plots and metrics
        ↓
save the trained model
        ↓
run future inference
        ↓
produce future_predictions.csv

The same model can also be run through the normal Python project entry points.

Where I want to take this next

I do not think the next useful experiment is simply v10 with another LSTM layer.

The next comparison should keep exactly the same evaluation framework and introduce genuinely different forecasting architectures.

Candidates include models such as PatchTST and N-HiTS, together with strong classical baselines such as regularized linear models.

The interesting question is no longer:

Can a neural network generate something that looks like tomorrow's stock price?

It is:

Can any of these models repeatedly extract out-of-sample information from financial time series that survives comparison with simple baselines across assets and market regimes?

That is a much harder problem.

It is also a much more interesting one.

Final thoughts

This project started in 2020 as an experiment with TensorFlow and LSTM networks.

Six years later, the biggest improvement is not a new neural-network layer.

It is the evaluation framework around the network.

A sophisticated model losing honestly to a naïve baseline is more useful than a beautiful chart that hides the comparison.

In the recorded benchmark, v8 outperformed v7 on RMSE in 17 of 20 asset/year test slices, while both models still had negative mean skill against the naïve baseline.

v9 moves the target towards returns, direction and uncertainty.

Neither result justifies claiming that the market has been “solved”.

That is exactly the point.

The repository is now structured to make those failures visible, reproducible and useful.

And for a forecasting research project, I think that is a much better place to start.


Project: JordiCorbilla/stock-prediction-deep-neural-learning

v9 notebook: examples/notebooks/stock_prediction_lstm_v9.ipynb

Benchmark: 2026-09-26 v7-v8 walk-forward report

Comments

Popular Posts