sktime/tserve

Time Series serving (TServe): local inference server for foundation models

Python

8

78 commits

updated Sep 28, 2026

See the code

See what people are saying

README

TServe

Documentation · Quick start · Models · API
ProjectLicense Python PyPI
StatusTests Docs Docker

Time series serving for foundation models. TServe loads models such as Chronos, TimesFM, Moirai, TTM, and TiRex once, keeps them in memory, and answers forecast requests over HTTP.

Each model family ships its own package, input format, and loading code. sktime wraps them as forecasters with one common interface, and TServe runs those forecasters as a server behind a single request: a table of past values and a horizon. Trying another model means changing one field, not rewriting your pipeline.

  • Over 100 checkpoints. Each family has its own Docker tag or pip extra, for CPU or GPU. Catalog · Capabilities
  • Loaded once, kept warm. Weights download and load at startup, so each request pays only for inference.
  • JSON from anywhere. POST /predict works from curl or any language. Send a prediction
  • Native tables in Python. The Client takes a dict, pandas, polars, or pyarrow table and returns predictions in the same type.
  • A dashboard in the browser. GET / plots a forecast from a sample series or your own CSV. What you can do
  • Your own sktime models. Serve a configured forecaster, a saved .zip, or a craft spec next to the catalog models.

TServe demo: start the server, query /models and /predict, then forecast in the dashboard

TServe is a server you run on your own hardware, not a hosted API. How it works · Docker Hub

First forecast

Docker is the short path. This image can load Chronos Bolt, Chronos T5, TTM, and TimesFM 2.x. The first start downloads the weights you name.

docker run --rm -p 8000:8000 sktime/tserve:hub chronos_bolt ttm_r3

When the log prints the local URLs, the models are warm. Five days of sales, three steps ahead:

curl -s http://127.0.0.1:8000/predict -H "Content-Type: application/json" -d '{
  "past": {
    "timestamp": ["2024-01-01", "2024-01-02", "2024-01-03", "2024-01-04", "2024-01-05"],
    "sales": [120, 135, 128, 142, 138]
  },
  "fh": 3,
  "model": "chronos_bolt"
}'
{
  "predictions": {
    "timestamp": ["2024-01-06T00:00:00", "2024-01-07T00:00:00", "2024-01-08T00:00:00"],
    "sales": [139.96, 138.93, 138.26]
  },
  "quantiles": null,
  "model": "chronos_bolt",
  "request_id": "…"
}

Open http://127.0.0.1:8000/, pick chronos_bolt, and plot the same series. The page can also take a pasted or dropped CSV. What you can do

The same call from Python. The client posts Arrow, and predictions comes back as the same kind of table you sent:

pip install "tserve[client]"
from tserve.client import Client

past = {
    "timestamp": ["2024-01-01", "2024-01-02", "2024-01-03", "2024-01-04", "2024-01-05"],
    "sales": [120, 135, 128, 142, 138],
}

with Client("http://127.0.0.1:8000") as client:
    result = client.predict(past=past, fh=3, model="chronos_bolt")

print(result.predictions)

The walkthrough, including GET /models and PowerShell: Quick start. A GPU host adds --gpus all and uses sktime/tserve:hub-gpu. GPU images

Dashboard

The running server serves a browser console at GET /. The model list is whatever this process loaded. You set a horizon, optionally a prediction interval, and a series (a built-in sample, pasted CSV, or a dropped file, parsed in the browser), then the page posts POST /predict and plots the result. Health and runtime stats sit on the right. What you can do

Below, timesfm_3 forecasts retail sales with 90% prediction interval.

TServe dashboard: timesfm_3 with a 90% prediction interval

Models

117 checkpoints. The extra name is the image tag, sktime/tserve:<tag>, and server publishes as :base. GPU tags append -gpu. base has no GPU tag. added counts checkpoints that extra contributes. full is the total, including naive.

naive always loads, so you can try the process before any download. GET /models lists what this process loaded, which is smaller than the catalog. What gets loaded

extrafamiliesaddedexample
serverNaive1naive
hubChronos Bolt, Chronos T5, TTM, TimesFM 2.x81chronos_bolt
chronosChronos-23chronos_2
kronosKronos, WindFM5kronos
graniteFlowState2flowstate
moiraiMoirai 2, Moirai 1.x, Lag-Llama8moirai_2
tirexTiRex2tirex
tirex2TiRex-24tirex_2
totoToto-25toto_2_0_4m
mantisMantis3mantis_8m
timesfm3TimesFM 31timesfm_3
t0T01t0
tafsutTafsut1tafsut
fullall of the above117chronos_2

kronos is built on base. Chronos Bolt, TTM, and TimesFM 2.x load on the images that include hub: chronos, granite, moirai, tirex, tirex2, toto, mantis, timesfm3, t0, tafsut, and full. TimesFM 3, TiRex-2, T0, and Tafsut load on their own extras and on full. Tags, GPU variants, and how the extras stack: Dependencies.

Each family page has its own start command. The catalog collects them under Start a server. Switching images is the tag plus the example from that row:

docker run --rm -p 8000:8000 sktime/tserve:moirai moirai_2

Multivariate series, covariates, and quantiles differ by family: Capabilities. Every checkpoint name: All models. mantis needs more than 127 rows of past: mantis.

Install

Docker needs no local Python. uv and pip need Python 3.12 or newer. Install the extra, or pull the tag, for the family in the table above.

Load a model

A bare tserve loads naive only. Name the models you want beside it. Flags are --host, --port, and --log-level: Flags · Startup and exit.

Startup prints the dashboard, Swagger, and ReDoc. What you can do · Live OpenAPI

Send a forecast

JSON goes to POST /predict. The Python client posts Arrow to POST /predict/bytes. Both send the same fields. Request fields

past is one row per timestamp, fh is how many steps ahead, and the forecast continues from the last row. Omit time and the first column is time. Omit target and the other columns are the series, except any you also put in future. Column roles · Column inference · Prediction horizon and model

you wantread
JSON from any languageSend a prediction · Endpoints
Row-oriented JSON, or ArrowUse row-oriented JSON · Arrow endpoint · Table formats
pandas, polars, or pyarrowUse native tables · Connect
A pandas DatetimeIndexUse an indexed pandas frame · Time
Covariates or a static rowFuture and static data · Request covariates
QuantilesQuantiles · HTTP · Python
The response shapeResponse
Health, loaded models, latencyInspect the server · Status routes

A body the schema rejects is 422. An unloaded model or a missing column is 400. Predict requests · Python client errors · Startup

Which families can take more than one target, a covariate, or a quantile: Capabilities. Panel and hierarchical input are outside this contract. Validation and limits

The generated reference for the same surface: HTTP API · POST /predict · Python API.

License

BSD 3-Clause. See LICENSE.

License covers only the model server, not the models themselves or distributions pathways such as Hugging Face. Third party model weights, model code, or distribution pathways may create their own implications via licenses or T&C. While we try to make it easy for users to gain a transparent picture of legal implications, we do not assume any liability or guarantee correctness of metadata related to third party licenses or T&C.

Development setup, checks, tests, and image builds: Development · Checks · Tests · Docker images.

deep-learning
docker
forecasting
foundation-models
gpu
huggingface
inference
machine-learning
model-serving
pretrained-models
python
pytorch
serving
sktime
time-series
timeseries
time-series-forecasting
transformers

sktime/tserve

Time Series serving (TServe): local inference server for foundation models

Python

8

78 commits

updated Sep 28, 2026

See the code

See what people are saying

README

TServe

Documentation · Quick start · Models · API
ProjectLicense Python PyPI
StatusTests Docs Docker

Time series serving for foundation models. TServe loads models such as Chronos, TimesFM, Moirai, TTM, and TiRex once, keeps them in memory, and answers forecast requests over HTTP.

Each model family ships its own package, input format, and loading code. sktime wraps them as forecasters with one common interface, and TServe runs those forecasters as a server behind a single request: a table of past values and a horizon. Trying another model means changing one field, not rewriting your pipeline.

  • Over 100 checkpoints. Each family has its own Docker tag or pip extra, for CPU or GPU. Catalog · Capabilities
  • Loaded once, kept warm. Weights download and load at startup, so each request pays only for inference.
  • JSON from anywhere. POST /predict works from curl or any language. Send a prediction
  • Native tables in Python. The Client takes a dict, pandas, polars, or pyarrow table and returns predictions in the same type.
  • A dashboard in the browser. GET / plots a forecast from a sample series or your own CSV. What you can do
  • Your own sktime models. Serve a configured forecaster, a saved .zip, or a craft spec next to the catalog models.

TServe demo: start the server, query /models and /predict, then forecast in the dashboard

TServe is a server you run on your own hardware, not a hosted API. How it works · Docker Hub

First forecast

Docker is the short path. This image can load Chronos Bolt, Chronos T5, TTM, and TimesFM 2.x. The first start downloads the weights you name.

docker run --rm -p 8000:8000 sktime/tserve:hub chronos_bolt ttm_r3

When the log prints the local URLs, the models are warm. Five days of sales, three steps ahead:

curl -s http://127.0.0.1:8000/predict -H "Content-Type: application/json" -d '{
  "past": {
    "timestamp": ["2024-01-01", "2024-01-02", "2024-01-03", "2024-01-04", "2024-01-05"],
    "sales": [120, 135, 128, 142, 138]
  },
  "fh": 3,
  "model": "chronos_bolt"
}'
{
  "predictions": {
    "timestamp": ["2024-01-06T00:00:00", "2024-01-07T00:00:00", "2024-01-08T00:00:00"],
    "sales": [139.96, 138.93, 138.26]
  },
  "quantiles": null,
  "model": "chronos_bolt",
  "request_id": "…"
}

Open http://127.0.0.1:8000/, pick chronos_bolt, and plot the same series. The page can also take a pasted or dropped CSV. What you can do

The same call from Python. The client posts Arrow, and predictions comes back as the same kind of table you sent:

pip install "tserve[client]"
from tserve.client import Client

past = {
    "timestamp": ["2024-01-01", "2024-01-02", "2024-01-03", "2024-01-04", "2024-01-05"],
    "sales": [120, 135, 128, 142, 138],
}

with Client("http://127.0.0.1:8000") as client:
    result = client.predict(past=past, fh=3, model="chronos_bolt")

print(result.predictions)

The walkthrough, including GET /models and PowerShell: Quick start. A GPU host adds --gpus all and uses sktime/tserve:hub-gpu. GPU images

Dashboard

The running server serves a browser console at GET /. The model list is whatever this process loaded. You set a horizon, optionally a prediction interval, and a series (a built-in sample, pasted CSV, or a dropped file, parsed in the browser), then the page posts POST /predict and plots the result. Health and runtime stats sit on the right. What you can do

Below, timesfm_3 forecasts retail sales with 90% prediction interval.

TServe dashboard: timesfm_3 with a 90% prediction interval

Models

117 checkpoints. The extra name is the image tag, sktime/tserve:<tag>, and server publishes as :base. GPU tags append -gpu. base has no GPU tag. added counts checkpoints that extra contributes. full is the total, including naive.

naive always loads, so you can try the process before any download. GET /models lists what this process loaded, which is smaller than the catalog. What gets loaded

extrafamiliesaddedexample
serverNaive1naive
hubChronos Bolt, Chronos T5, TTM, TimesFM 2.x81chronos_bolt
chronosChronos-23chronos_2
kronosKronos, WindFM5kronos
graniteFlowState2flowstate
moiraiMoirai 2, Moirai 1.x, Lag-Llama8moirai_2
tirexTiRex2tirex
tirex2TiRex-24tirex_2
totoToto-25toto_2_0_4m
mantisMantis3mantis_8m
timesfm3TimesFM 31timesfm_3
t0T01t0
tafsutTafsut1tafsut
fullall of the above117chronos_2

kronos is built on base. Chronos Bolt, TTM, and TimesFM 2.x load on the images that include hub: chronos, granite, moirai, tirex, tirex2, toto, mantis, timesfm3, t0, tafsut, and full. TimesFM 3, TiRex-2, T0, and Tafsut load on their own extras and on full. Tags, GPU variants, and how the extras stack: Dependencies.

Each family page has its own start command. The catalog collects them under Start a server. Switching images is the tag plus the example from that row:

docker run --rm -p 8000:8000 sktime/tserve:moirai moirai_2

Multivariate series, covariates, and quantiles differ by family: Capabilities. Every checkpoint name: All models. mantis needs more than 127 rows of past: mantis.

Install

Docker needs no local Python. uv and pip need Python 3.12 or newer. Install the extra, or pull the tag, for the family in the table above.

Load a model

A bare tserve loads naive only. Name the models you want beside it. Flags are --host, --port, and --log-level: Flags · Startup and exit.

Startup prints the dashboard, Swagger, and ReDoc. What you can do · Live OpenAPI

Send a forecast

JSON goes to POST /predict. The Python client posts Arrow to POST /predict/bytes. Both send the same fields. Request fields

past is one row per timestamp, fh is how many steps ahead, and the forecast continues from the last row. Omit time and the first column is time. Omit target and the other columns are the series, except any you also put in future. Column roles · Column inference · Prediction horizon and model

you wantread
JSON from any languageSend a prediction · Endpoints
Row-oriented JSON, or ArrowUse row-oriented JSON · Arrow endpoint · Table formats
pandas, polars, or pyarrowUse native tables · Connect
A pandas DatetimeIndexUse an indexed pandas frame · Time
Covariates or a static rowFuture and static data · Request covariates
QuantilesQuantiles · HTTP · Python
The response shapeResponse
Health, loaded models, latencyInspect the server · Status routes

A body the schema rejects is 422. An unloaded model or a missing column is 400. Predict requests · Python client errors · Startup

Which families can take more than one target, a covariate, or a quantile: Capabilities. Panel and hierarchical input are outside this contract. Validation and limits

The generated reference for the same surface: HTTP API · POST /predict · Python API.

License

BSD 3-Clause. See LICENSE.

License covers only the model server, not the models themselves or distributions pathways such as Hugging Face. Third party model weights, model code, or distribution pathways may create their own implications via licenses or T&C. While we try to make it easy for users to gain a transparent picture of legal implications, we do not assume any liability or guarantee correctness of metadata related to third party licenses or T&C.

Development setup, checks, tests, and image builds: Development · Checks · Tests · Docker images.

deep-learning
docker
forecasting
foundation-models
gpu
huggingface
inference
machine-learning
model-serving
pretrained-models
python
pytorch
serving
sktime
time-series
timeseries
time-series-forecasting
transformers

Languages

Python

73.3%

JavaScript

11.3%

CSS

9.1%

HTML

4.0%

HCL

1.5%